The World Model Is Necessary but is it enough
Embodied AI needs more than predictive prowess; it requires institutional permission to act on predictions for real-world impact.
Michael Carroll | The One-Degree Dispatch | Industrial AI and Enterprise Agency The World Model Is Necessary. Is It Enough. World models may give AI the ability to predict and plan in the real world. Enterprise agency begins only when prediction can cross permission, audit, and decision rights. By Michael Carroll Founder | Investor | Research Fellow | Board Advisor | Industrial AI, Causal Systems, and Enterprise Transformation
Lead image. The visible failure is only the beginning. The harder question is who, or what, has the authority to act before the loss reaches the ledger. On January 22, 2026, the room at AI House Davos was too small for the size of the claim being made. The audience was packed close enough that the logistics became part of the record, with a later comment asking why roughly fifty people were put in a room that looked built for a dozen. A glass sat on the table, heavier than the hand expected. A microphone squealed for a second, then the session settled into the kind of technical confidence that can make a hard problem feel almost solved. The subject was embodied AI, the class of systems that must see, hear, predict, and act in the world rather than merely produce language about it. The argument drew a sharp boundary between text and reality. Language is already cut into tokens that sit near meaning, while the physical world arrives as continuous, high-dimensional, noisy signal. Video was the obvious example, but the same burden applies to a robot, a turbofan, a warehouse, a mill, a transmission grid, or a production process where state changes faster than a human can inspect every variable. The session's most useful claim was not that progress had stalled. It was that much of the market had mistaken fluent output for understanding. A system that speaks well has not proved that it understands what will happen if it acts. Consumer robots, reliable autonomy, and practical embodied agents will require internal representations that predict the consequences of action and plan across time. The audience had reason to nod because that diagnosis is mostly right. The omission sits after the diagnosis, not inside it. A machine may learn to predict what action will do to the world. A firm still has to decide who, or what, has the right to act on that prediction when money, safety, reputation, or liability is on the line. That is where technical agency becomes institutional agency, and that is where many expensive systems stop moving. A world model that cannot cross permission is still trapped inside the recommendation layer.
Prediction Does Not Move the Enterprise
The strongest version of the world model thesis is clean. Given the state of an environment at time t, and given an action, can the system predict the state at time t plus one. If it can make that prediction in an abstract representation rather than raw pixels, it can plan without drowning in irrelevant detail. If it can stack those models across different time horizons, it can plan hierarchically. If it can plan, then it may begin to act. That ladder matters because it explains why many current agent systems look better in demos than in the open world. They follow scripts. They call tools. They repeat sequences that worked before. They appear competent until the task is new, the environment changes, or the next step depends on a consequence the system did not predict. Some humanoid demonstrations are impressive as engineering. They are not proof of general intelligence when the behavior has been precomputed or tightly choreographed. For robotics, the missing piece can look technical. The robot must connect perception to control. It must know enough about its own body and the objects around it to choose an action that will not fail. In that setting, authority is often assumed because the robot owns its actuators. When the system plans to move an arm, the arm either moves or it does not. The gap between deciding and doing is short. In enterprise operations, that gap is the system. A model can detect a failure pattern in milliseconds, yet the line may keep running because no one has authority to stop it without a sign-off. A process may show drift before the customer feels it, yet the organization waits for more evidence because the next move will show up in cost, service, inventory, or bonus calculations. A maintenance signal may be right, but the work order still competes with production commitments and a supervisor's fear of being the person who called the failure too early. The firm is also predicting. It predicts blame. It predicts the meeting after the event. It predicts whether the incident review will punish a false alarm more than a late call. It predicts whether the person who acted will be protected if the model was wrong. Those predictions are not irrational. They are the reason competent people often choose delay when the spreadsheet says delay is costly. A world model that ignores that institutional environment is not planning for the world in which the action must occur. It is planning for a cleaned-up version of reality where inference automatically becomes permission. That world does not exist in large firms. The physical process may allow action in seconds. The organization may require a meeting, a policy check, a customer approval, a legal review, or an escalation path that takes longer than the window in which the decision still matters. The False Comfort Inside a Correct Technical Claim The world model argument is persuasive because it rejects a cheap answer. It refuses to call text generation a substitute for common sense. It refuses to treat imitation as open-world competence. It insists that agents need predictive structure, abstraction, and planning. After several years of hearing the word agent attached to little more than tool chaining, that discipline feels overdue. The same discipline can become a new false comfort. In technical circles, it sounds like a research program. Build the representation. Add planning. Connect the planner to action. In boardrooms, the same idea becomes a capital allocation story. Fund the perception stack, buy the planning company, install the model, and wait for operating speed to improve. The reason the belief survives is that it contains real truth. Without prediction, there is no reliable action. What would have to be true for this outcome to keep repeating. The repeating outcome in industrial firms is not ignorance. It is failed conversion. Companies see more than they once saw and still act late. They predict more and still absorb the cost. They detect anomalies earlier and still lose margin after the fact. Visibility has improved faster than permission, which means the enterprise receives earlier warnings while preserving the same old bottleneck between signal and action. What is observed is a serious technical argument that world models are central to embodied AI. What is inferred is that executives will import that argument into enterprise autonomy as if capability were the binding constraint. What is projected is that many deployments will disappoint when permission, audit, and decision rights are treated as downstream integration work rather than part of the architecture itself. This is the point at which the claim has to be pressed. A world model may predict consequences. It does not grant legitimacy. A planner may choose an action. It does not decide whether the action is authorized. A system may be right about the physics and wrong about the firm, because the firm is not only a physical environment. It is a legal, financial, political, and behavioral one. The mistake is familiar. Companies have bought dashboards that made the truth more visible without changing the speed at which decisions were made. They built control rooms that improved awareness while leaving authority in the same old chain. They added analytics to reviews and then wondered why the review still consumed the week. The screen changed. The decision right did not.
Supporting image. When the approval chain is slower than the operating window, the model may be right and still be powerless.
Figure 1. From Signal to Authority. The missing step is not more visibility. It is the controlled conversion of prediction into legitimate action. The Enterprise Actuator Is a Decision Right In robotics, embodiment means sensors, actuators, motion, physical state, and feedback. That definition is right for a machine that must move through the physical world. For industrial autonomy, embodiment also means the system is inserted into an enterprise with policies, incentives, approval chains, work practices, contracts, and liability. Those constraints are not softer than physics. They determine what action can be taken, who can take it, when it can be taken, and who owns the bill when it goes wrong. A model that does not include those constraints is partial by design. It may know that a process is likely to fail if left alone. It may know that intervention now is cheaper than intervention later. It may even choose the right intervention. Yet the decision can still stall because the model has reached the boundary where the enterprise asks for something prediction alone cannot supply. It asks for a defensible reason to allow action. That reason is not a prettier explanation in natural language. Explanation can become another theater of persuasion, especially when a confident model turns statistical association into a clean story. The best story may win the meeting without being the truest account of cause. Leaders know this because they have watched human organizations do the same thing for decades. Under pressure, stories become shields. Trust at scale is different. It is not a feeling created by eloquence. It is a property of a system that can be checked after the fact. When a decision has material consequence, the organization needs to reconstruct what the system believed, what evidence supported that belief, what assumptions were active, what policy constraints applied, what authority tier was used, what action was taken, what happened next, and what the system learned. That is not paperwork. It is how delegated authority survives contact with accountability. The enterprise actuator is not a motor. It is a decision right.
Figure 2. Why Intelligence Stalls in Advisory Mode. The failure is not that the machine failed to see. The failure is that the enterprise preserved the old permission path. This is where audit enters the argument not as compliance overhead, but as operating machinery. If the chain from signal to action cannot be reconstructed, the machine will remain advisory for every decision that matters. It will recommend. It will score. It will surface risk. Then a human will step between the system and the action because the organization cannot defend letting the system cross the line on its own. That human intermediation will often be described as prudence, and sometimes it will be prudence. Some actions should remain human-owned. Some contexts should require consent. Some evidence should trigger escalation rather than automation. The hard fact is that firms rarely distinguish those cases cleanly before deployment. They buy autonomy with the language of speed and then govern it with the habits of manual review.
Supporting image. Audit is not ceremony when the action has consequence. It is the record that lets authority be delegated without pretending risk disappeared. Restricted Domains Are Not a Weakness. They Are the Proof The fair counterargument begins with the strongest version of the technical case. Open-world action cannot be built on scripts alone. A robot that only repeats sequences cannot handle the world when the sequence breaks. A vehicle that merely imitates human driving from video will struggle with rare events, broken assumptions, and the long tail of physical risk. A system that predicts pixels may waste effort reconstructing what does not matter and miss the abstract state that does. That argument should be taken seriously. It is why the world model thesis has force. If a system cannot predict the likely consequences of its own actions, it should not be called an agent in any serious sense. It can automate a task. It can assist a user. It can generate a plan-shaped artifact. It cannot reliably shape an outcome when the world stops following the script. There is still a counterexample that cannot be waved away. Scripted systems do create value in bounded settings. Level four autonomy can be practical inside restricted operating domains. Expert systems failed as a general road to intelligence, but they did not fail everywhere. Rules, maps, operating envelopes, sensors, and constraints have produced real performance precisely because they limited the conditions under which action could occur. That lesson should not embarrass the autonomy argument. It should refine it. Restricting the domain is how a system earns the right to act before it earns the right to generalize. Enterprises already do this with people. A technician may authorize one action, a plant manager another, a CFO another, and the board another. Authority expands with consequence, evidence, role, history, reversibility, and trust. The same ladder has to exist for machines. Special purpose intelligence is not weaker product language. It is the practical form of responsibility. A bounded intelligence with a bounded action envelope can be audited, measured, corrected, and trusted with a narrow class of decisions. A general system that claims it can do anything cannot be granted broad authority because no one can define the liability boundary around everything. Bounded autonomy scales. Unbounded autonomy gets meetings.
Figure 3. The Permission Ladder. The enterprise can delegate narrow action before it delegates broad judgment. This does not contradict the world model thesis. It completes it. The world model expands what the system can understand and plan. The permission ladder defines what the system is allowed to do with that understanding. One is the technical substrate. The other is the deployment path. Confusing the two will make capability look like progress until the first expensive decision has to be explained. The proof will not arrive in a laboratory benchmark. It will arrive in the exception queue, the maintenance hold, the customer allocation call, the quality containment decision, the supplier substitution, the energy curtailment, and the production stop. Those are the places where intelligence has to become motion while someone still owns the consequence. If the system can act only when nothing material is at stake, the enterprise has not gained agency. It has gained a faster recommender. Causal Governance Is the Missing Tier The administrative object that matters most may be the audit record. Not the dashboard. Not the model card. Not the post-event slide. The record. What did the system believe when it acted. What mechanism did it claim. What intervention did it select. What evidence crossed the threshold. Which policy allowed the move. Which human retained override. What result came back. What changed afterward. A world model predicts. A scientist tests. The difference is decisive because enterprises do not merely need a prediction that X will lead to Y. They need a defensible mechanism for why X is believed to cause Y under the current conditions, and they need a record of whether intervention changed the outcome. That is why causal governance cannot be added as a decorative layer after the model is deployed. It is the layer that makes delegated action defensible. The Automated Scientist becomes important for that reason. It is not a branding flourish or a research metaphor. It is the executive function that tests mechanism claims, tracks interventions, records assumptions, separates correlation from cause where possible, and updates what the enterprise can responsibly make true. It does not replace judgment. It gives judgment a record strong enough to survive scrutiny when the model is right, when it is wrong, and when the evidence is mixed. This is also why language alone cannot carry trust. A fluent explanation may help a person understand what the system appears to be doing. It does not prove that the system had the right evidence, applied the right policy, respected the right boundary, or learned the right lesson. In high-consequence settings, trust has to be rebuilt as a chain of reason that can be inspected. If that chain is missing, the safest-looking governance answer will be to keep the system behind glass. A CEO should ask where inference becomes permission inside the firm. Is the state change created by evidence, by role, by politics, or by fear. When a system raises a high-confidence risk signal, what is the fastest action the organization will allow without escalation. What is the slowest action it routinely forces into a meeting. Does that boundary map to financial materiality, customer harm, safety risk, and reversibility, or does it map to the identities of the people who might get blamed. A CFO should ask a harsher version. If the model is wrong and acts, who owns the bill. If the model is right and cannot act, who owns that bill. Which failure does the company punish more in practice. The answer will reveal whether the firm is designed to reduce loss or to protect explanations after loss has already arrived. Audit is not paperwork. It is the price of delegated authority.
Figure 4. The Complete Enterprise Agency Stack. Prediction needs governance, audit, boundaries, and decision rights before it becomes agency. Those questions make the thesis testable. If world models arrive with causal audit, permission tiers, and bounded authority, industrial autonomy will convert faster than many cautious observers expect. If they arrive as better prediction engines wrapped in natural language explanations, they will hit the same wall that visibility programs hit. The system will see the risk earlier, and the organization will still have to negotiate the right to touch it. The Failure Will Look Technical Until the Bill Is Read Here is the prediction that would be costly to get wrong. By May 2028, many industrial firms will have deployed planning-capable models that can identify failure modes earlier than their current systems. Most of those deployments will remain advisory for actions that are financially, operationally, or politically material, unless the firms build permission ladders and causal audit ledgers at the same time. The disappointment will be described as a model limitation. In many cases, it will be an authority limitation. The prediction would be weakened if firms begin changing decision rights as fast as they buy planning capability. It would be weakened if audit-grade causal records become a standard feature of industrial AI deployments rather than a later governance project. It would be weakened if executives stop treating autonomy as a software purchase and begin treating it as an operating model decision. Those conditions are possible. They are not yet the default behavior of large firms. The cost of getting this wrong is not limited to wasted software spend. Earlier warning without authority often creates more anxiety inside the organization because the firm now knows sooner what it still cannot move to prevent. Operators watch risk accumulate. Managers ask for more proof. Executives see the same charts more often. The company feels more informed while the avoidable loss keeps moving through the system. Earlier warning without authority often becomes earlier anxiety.
This is why the next serious autonomy conversation has to include the operating system of the firm, not just the model architecture. The issue is not whether world models are necessary. They are. The issue is whether the enterprise is willing to rebuild the path from signal to authority so that prediction can become responsible action. Without that path, the model will sit in the familiar posture of enterprise technology. Useful, impressive, expensive, and trapped. The world model may become the right technical spine for embodied AI. It may help machines represent the world at the level where planning becomes possible. It may make robots, vehicles, and industrial systems far more capable than the current generation of brittle scripts and fluent assistants. None of that gives the machine the right to act in an enterprise when consequence is material. Rights of action are designed, granted, bounded, audited, and revoked. That is the hard distinction. Intelligence can be built in the model. Agency has to be granted by the system around it. The first problem belongs to research and engineering. The second belongs to leadership, governance, finance, operations, and law. A firm that solves only the first will have better predictions. A firm that solves both may have a new form of control. The small room at Davos captured the right technical turn. The old language-first story is no longer enough for the world that has to be acted upon. But the enterprise lesson is sterner. The machine can see the failure, model the consequence, and select the better move. Until the firm has decided how that move becomes legitimate, the future remains visible and out of reach. The world model is necessary. It is not enough. References This article draws on AI House Davos 2026 and its January session titled Embodied AI: Systems that See, Hear, and Act in the World Alongside Humans, featuring Yann LeCun and Marc Pollefeys, as well as LeCun's 2022 paper A Path Towards Autonomous Machine Intelligence for the technical case behind world models, abstract representation, and planning. It uses the 2025 research survey Embodied AI Agents: Modeling the World and related embodied AI survey work to ground the distinction between language systems, world models, planning, and action in physical environments. The argument about prediction versus mechanism is anchored in Judea Pearl's Causality and The Book of Why, while the operating logic of systems, variation, and management responsibility draws on W. Edwards Deming's Out of the Crisis. The authority and coordination argument draws on Ronald Coase's 1937 theory of the firm, Herbert Simon's Administrative Behavior and bounded rationality, Daniel Kahneman's Thinking, Fast and Slow for the cost of deliberation and automatic judgment, and Jensen and Meckling's 1976 agency-cost work for why delegated authority requires monitoring, incentives, and audit. The enterprise claim is also informed by practical governance lessons from industrial quality, safety, maintenance, and operations, where the costliest failures often occur after the signal was visible but before action was authorized.