The One-Degree Dispatch

The Copilot Question We Asked in February 2024 (Draft)

2024 · The Nature of Intelligence · 3,409 words

The essay warns that overreliance on analytics without actionable expertise can exacerbate performance gaps in manufacturing.

https://www.linkedin.com/pulse/looming-copilot-disaster-manufacturing-navigating-complexcarroll-qs If you want to understand the Copilot moment in manufacturing, you have to start there. Not with a product launch. Not with a demo. Not with a vendor story. You have to start with the uncomfortable reality that information has been rising for years while outcomes have not moved in proportion. That gap is not a technology gap. It is a conversion gap. In that February 2024 conversation, Paul and I drifted from the chart into the machinery beneath it. We talked about how analytics, used the way it is often used, keeps organizations in the shallow end of the solution pool. We curve-fit symptoms. We polish charts. We give sophisticated names to the act of describing what we already know is wrong. Then we act surprised when the organization gets better at explanation but not better at performance. Then we reached the hinge. The part that now reads less like commentary and more like a warning label. We asked whether the obsession with new analytics had become a distraction at the exact moment industry was losing the people who could turn analytics into action. Deep tenure. Disciplined operations knowledge. Frontline knowhow. The adaptive capacity to look at noise, find signal, and act under constraint. And then we asked the question that now defines the Copilot era. What happens if we respond to a knowledge collapse by giving people more to think about. That was the core concern in 2024. Not that generative AI was “bad.” Not that copilots were a gimmick. The concern was that we were asking humans to do more inference at the exact moment we had fewer humans left who could do inference well. Since then, that concern has become sharper, because we can now see what we could only sense then. We have learned that copilots can be genuinely useful in the right context. We have also learned that usefulness is not the same thing as value. We have learned that time saved is not the same thing as productivity gained. We have learned that adoption is not a straight line when trust, integration, and accountability are real. We have learned, above all, that the enterprise is not short on intelligence. It is short on architectures that can convert evidence into action. So let’s return to that February 2024 article with professional discipline. Not with bravado. Not with self-congratulation. With the only posture that earns trust. Precision. Were Paul and I right. Where were we wrong. What did the evidence teach us. I keep returning to a scene that has played out in a thousand plants, quietly, without headlines.

A supervisor sits in front of three screens that all claim to be the truth. One system says the line is running. Another says the work order is still open. A third says quality is holding material the dock is already staging. The historian trace shows a drift that looks harmless until you know what it does to viscosity, then to yield, then to scrap, then to the customer complaint that arrives two weeks later after the root cause has been buried under the next fire. An operator is waiting. A maintenance tech is waiting. A planner is waiting. The organization is waiting, because the organization is built to wait until the right person has interpreted the signal, assembled context, and secured permission. This is manufacturing. Evidence becomes consequence in real time. The enterprise is judged not by what it knows, but by what it can convert into action. The winners are not the companies with the prettiest dashboards. The winners are the companies whose decision loops move faster without breaking trust. That is why the Copilot story has been misunderstood. Copilots were sold as tools for individual productivity. Manufacturing is not an individual productivity contest. It is system behavior. It is throughput of the constraint. It is stability of the line. It is containment of deviation. It is prevention of escape. You can save minutes in email and still lose hours in decision latency. You can draft faster and still stabilize slower. You can be “productive” in the office layer and still be ineffective in the operational layer. Paul’s productivity question was never an attack on analytics. It was an accusation against a lazy assumption. The assumption that tools bend outcomes by existing. That assumption walked straight into the Copilot era. In 2024, Paul and I began to question the validity of requiring humans to do inference and giving them more to think about at the exact moment the number of people in industry with deep tenure and adaptive capacity was collapsing. The few who remained were also being asked to oversee a dramatically expanded surface area of systems, alerts, dashboards, exceptions, and “marginal improvements” justified one capital line item at a time. This is the part most leaders do not see because it hides inside reasonable decisions. Each addition is defensible on its own. A sensor justified by a local ROI. A dashboard justified by a department’s metrics. A governance step justified by a compliance audit. Another tool justified by a perceived gap. In aggregate, those rational moves create a cognitive tax. The enterprise becomes a machine that produces more signals than its people can safely interpret. Then it blames the people when they cannot keep up.

So when we say “it feels like 8 times more,” we are not doing drama. We are describing compounding. Half as many true experts. More surface area. More reconciliation. More handoffs. More meetings. More queues. More exceptions. More permission steps. More context-switching. More verification. More fatigue. Then we put a copilot next to the person and called it relief. If the copilot is menu-driven in practice, and if it requires the human to supply context, interpret output, validate correctness, and chase permission across systems, then it is not removing burden. It is consuming attention. It is the operational equivalent of texting and driving. Not because people are careless. Because the system is asking a mind to do too much at the moment it must control reality. That was the fear in 2024. It is now the fault line. So where were we right. We were right to focus on uneven impact, and we were right to treat it as a warning label, not a curiosity. In the February 2024 article, I anchored on an experiment by Otis and colleagues. 640 entrepreneurs in Kenya were given access to a GPT-4 powered mentor via WhatsApp. The result was not a clean average lift. High performers improved. Lower performers declined. The impact was uneven. The most important part was not the headline. It was the mechanism. The authors report that the divergence did not appear to come from different advice being given, but from how people selected and implemented advice. That mechanism matters in manufacturing because manufacturing is selection and implementation under constraint. A plausible answer is not valuable if it is implemented in the wrong order. A confident answer is not valuable if it violates local process physics. A fast answer is not valuable if it triggers a safety event, a quality escape, or a cascade of rework. So yes, we were right that copilots amplify capability. They do not equalize it by default. They are not a universal multiplier. But we were also wrong, or at least incomplete, and this is where discipline matters more than ego. The broader evidence since then shows that the direction of uneven impact depends on the structure of work.

In customer support settings, Brynjolfsson, Li, and Raymond found generative AI assistance increased productivity on average, with some of the largest gains accruing to less experienced workers in that context. So the correct statement is not, “AI helps the best and harms the rest.” The correct statement is sharper. AI amplifies the environment. When work is bounded, feedback is rapid, and correctness is checkable, AI can compress the learning curve. When work is open-ended, feedback is delayed, and quality depends on judgment and implementation, AI can widen dispersion and penalize weak inference. Manufacturing includes both regimes. A deviation response is not a drafting task. A maintenance triage is not a meeting summary. A quality containment decision is not an email rewrite. In 2024, we captured the risk. We did not fully articulate the conditionality. That is the first correction. The second thing we were right about was the conversion problem. Time saved is not productivity gained. This is where the last two years have been clarifying, because public evaluations have been forced to admit what most vendor stories avoid. The UK government ran a large Microsoft 365 Copilot experiment and reported an average time saving of about 26 minutes per day, along with broadly positive sentiment. That is real. It matters. But it does not settle the question manufacturing leaders care about. The Department for Business and Trade evaluation, which tries to account for overhead in ways most public evaluations do not, did not find evidence that time savings translated into improved productivity at the departmental level. This is the missing middle that defines the Copilot era. Minutes saved do not automatically turn into outcomes improved. A system can become locally more efficient and globally unchanged. A person can work faster and the enterprise can still be slow, because the bottleneck is not typing speed. The bottleneck is decision flow under permission. Manufacturing does not win because it drafts better. Manufacturing wins because it stabilizes faster. Manufacturing wins because containment happens earlier. Manufacturing wins because downtime is prevented rather than narrated.

In 2024, Paul’s question was already pointing at that. If fifteen years of analytics did not bend the productivity line, why would copilots do so automatically. The answer is they will not, unless the operating model changes. The third thing we were right about was verification, and the cost of verification is now impossible to ignore. If the output is not trusted, it must be verified. Verification takes time. Verification takes attention. Verification takes the scarce resource you are already losing when deep tenure collapses. The DBT evaluation tried to quantify this in a way most marketing material never will. It adjusted time savings downward when Copilot outputs were not used, and it treated “novel tasks” induced by Copilot as negative time savings because those tasks would not have existed otherwise. That is not an academic nuance. That is the heart of why copilots stall in high-consequence domains. A probabilistic system that generates plausible text can be extremely useful. But in a plant, plausible is not a standard. Safe is the standard. Correct is the standard. Verified is the standard. So verification becomes the second shift. If leaders do not count that cost, they are not measuring ROI. They are writing a story. The fourth thing we were right about was access and integration, and the market has started to surface that reality publicly. Reuters reported in late 2025 that Microsoft lowered sales growth targets for some AI products after sales staff missed goals, and it described customer resistance tied in part to practical integration challenges. One cited example involved reduced spending on Copilot Studio due to data integration issues. Ignore the headline and listen to the mechanism. Copilots promise broad usefulness. Enterprises are built on narrow permissions, fragmented systems, and real accountability. If the tool cannot reliably reach the evidence that matters inside the workflow that matters, it will disappoint. Not because the model is weak. Because the enterprise is not architected for conversion. This brings us to what we learned that we did not say clearly enough in 2024. The decisive constraint is not intelligence. It is permission.

Access is the architecture of power. If your systems cannot grant the right access safely, and if your workflow cannot turn evidence into permitted action, then copilots will remain assistance tools sitting beside a person who still must do inference, still must do reconciliation, still must do approval-chasing. That is not transformation. It is overlay. So where were we wrong. We were wrong to imply the “disaster” would show up as a dramatic event. The truth is that most operational failures are administrative before they are catastrophic. The enterprise drifts. The meetings multiply. The handoffs increase. The queues lengthen. The verification work grows. The high performers carry more weight. The system becomes brittle. You can have satisfied users and stagnant outcomes. You can have reported time savings and no measurable productivity conversion. The UK findings show positive sentiment and reported savings. The DBT evaluation found no evidence that those savings translated into productivity improvement at the departmental level. That is what “disaster” looks like in real enterprises. Drift, not drama. We were also wrong to treat “copilot” as a single category. Since 2024, the term has expanded to include embedded assistance, retrieval, automation, and early agent-like patterns. Microsoft Research has published results from a large randomized experiment suggesting measurable tasklevel benefits in certain knowledge work tasks for users who adopted it. So the honest position is this. Copilots can produce real local efficiency in language-heavy tasks. The question is whether the enterprise converts that local efficiency into system throughput and risk reduction. The question is whether copilots are deployed in the right task regimes, with the right governance, and with a permission architecture that reduces human intermediation instead of increasing it. That is the integrated learning. The Copilot era is forcing a more honest diagnosis of what actually constrains manufacturing performance. The constraint is not a lack of answers. The constraint is the fact that the enterprise keeps requiring humans to be the inference engine across an exploding surface area of complexity, under permission constraints that slow action. If you take that seriously, the strategic objective changes.

You stop buying answer engines and calling them intelligence. You start building inference removal. Inference removal means the enterprise stops requiring humans to be the integration layer between systems. It stops requiring humans to translate signal into action through meetings. It means permissions are modeled explicitly so that safe action can occur under constraints without re-requesting approval each time. This is not a philosophical preference. It is an operational necessity when deep tenure declines and complexity rises. And it leads to the only measurement discipline that can cut through the noise. Decision latency. Event. Detection. Reasoning. Intervention. Stabilization. If your AI deployment does not reduce that chain for a real operational loop, without increasing risk, then you did not buy intelligence. You bought activity. That is the line Paul’s productivity chart was trying to tell us in 2024. Visibility without conversion is entertainment. Copilots can accelerate visibility. They can accelerate narration. They can accelerate drafts. They cannot, by default, delete approvals, remove queues, or collapse decision latency. Not without architectural redesign. So yes, the February 2024 article feels prophetic now, but not because it predicted a backlash against a particular product. It feels prophetic because it described an architectural mismatch. We were about to repeat a familiar pattern. Add tools. Add surfaces. Add alerts. Add dashboards. Add AI. Then leave the decision geometry untouched. When you leave decision geometry untouched, the productivity line stays flat. The organization just becomes more articulate about why. That is the truth that the Copilot era is now forcing into the open. We should end where we began, with Paul’s question, because it is the same question hiding behind every GenAI spend decision today. Why did fifteen years of analytics not bend the line. Because analytics increased knowledge without changing authority. It increased visibility without deleting approvals. It increased awareness without reducing handoffs. It made the enterprise better at describing problems without being better at acting.

Copilots, deployed naively, will do the same. They will make the enterprise more fluent. They will not make it faster where it matters. They will not make it safer by default. They will not make it more resilient unless they are embedded in decision architectures that convert evidence into permitted action. So the real question is not whether copilots are useful. The real question is whether leadership will finally do the architecture work it has been postponing. Whether it will stop loading the mind while demanding higher-speed control. Whether it will stop rewarding marginal additions that create global complexity. Whether it will redesign decision loops so fewer humans must do inference, and the humans who remain can spend their judgment where judgment is truly required. That is what we now know. That is what we learned. And that is what we only suspected, but could not yet prove, on February 2, 2024.

References

Michael Carroll and Paul Boris. “The Looming CoPilot Disaster in Manufacturing? Navigating the Complex Landscape of Generative AI in Industry and Education.” February 2, 2024. Originating conversation with Paul Boris. Productivity chart framing. Demographic and knowledge erosion point. Initial thesis about inference burden and uneven GenAI impact. Otis, Clarke, Delecourt, Holtz, Koning. “The Uneven Impact of Generative AI on Entrepreneurial Performance.” Field experiment evidence of heterogeneous impacts and the role of selection and implementation of advice. Brynjolfsson, Li, Raymond. “Generative AI at Work.” Evidence of productivity effects in customer support settings, including heterogeneity and large gains for less experienced workers in that context. UK Government. “Microsoft 365 Copilot Experiment. Cross-Government Findings Report.” Reported time savings and user sentiment at scale. UK Department for Business and Trade. “Evaluation of the M365 Copilot Pilot.” Finds no evidence that time savings led to improved productivity. Accounts for overhead when outputs are unused or when Copilot induces additional tasks.

Microsoft Research. “Early Impacts of M365 Copilot.” Results from a large randomized experiment across firms showing measurable task-level effects and adoption variability. Reuters. Reporting on enterprise adoption friction and customer resistance signals, including integration challenges cited in Copilot Studio context. This argument draws on and adapts prior Chief Architect Network work rather than citing it verbatim, including Enough Intelligence. Shaping Destiny Without Digital Gods. The Line Between First Generation AI and Second Generation AI. The Question Engine. One Degree for Everyone and Everything. The Architecture of Permission No One Admits They Are Running.

LinkedIn Launch Post We were not arguing about AI in February 2024. We were arguing about architecture

In early 2024, Paul Boris asked me a question that still has not been answered by most GenAI deployments. If we have had fifteen years of analytics, dashboards, and “insights.” Why did productivity not move. That question started Paul and my February 2, 2024 article. It was a warning, not a prediction. Do not force humans to do more inference at the exact moment deep tenure is collapsing. Since then, the Copilot era arrived. So did the signal. Here is what we now know. Copilots can save time. They can also create a second shift called verification. They can make people feel faster while the enterprise stays slow, because decision latency and permission staircases remain untouched. Paul and I kept coming back to the same uncomfortable arithmetic. Fewer people with deep tenure and adaptive capacity. More surfaces to monitor. More alerts. More systems. More handoffs. More menus. More “bring your own logic.” Then we add a copilot and call it relief. That is not relief. That is texting and driving in an operating system built on cognitive overload. So we rewrote the argument with evidence. Not vendor claims. Not vibes. Evidence. What we were right about. Copilots amplify capability and risk depending on task regime. What we were wrong about. The failure mode is not dramatic. It is drift. What we learned. You do not win by generating more words. You win by removing inference burden and collapsing decision latency under explicit permissions.

If you are a COO, CIO, EHS leader, quality leader, plant manager, or anyone accountable for outcomes under constraint. This is your reading. Because Copilot is not the decision. Architecture is. Original article from Feb 2, 2024. The conversation that started it. https://www.linkedin.com/pulse/looming-copilot-disaster-manufacturing-navigating-complexcarroll-qseye/ New article. The updated evidence. The corrected thesis. PASTE NEW LINK HERE @Paul Boris. Thank you for asking the question that still exposes the bluff and co-authoring the articles. #Manufacturing #IndustrialAI #GenAI #MicrosoftCopilot #DigitalTransformation #OperationalExcellence #Safety #EHS #Quality #SupplyChain #CIO #COO

First Comment To Pin

Two links for anyone who wants the full arc. Original (Feb 2, 2024). https://www.linkedin.com/pulse/looming-copilot-disaster-manufacturingnavigating-complex-carroll-qseye/ Updated (2025). PASTE NEW LINK HERE If you are deploying copilots in operations, EHS, quality, maintenance, or supply chain. Read the updated one first.

Topics: synthetic-agency, causal-aiOpen in the Radiant ↗All dispatches