The One-Degree Dispatch

2025 Was the Year of Agent Washing

2025 · The Nature of Intelligence · 4,160 words

The shift from agent washing to outcome washing reveals a critical mismatch between software promises and actual deliverables, reshaping enterprise expectations and market trust.

asking it to carry more of the work, more of the consequence, and more of the distance between signal and action. What would have to be true for this outcome to keep repeating. That question sounds narrow until it is asked aloud in a budgeting meeting. Then it accuses a central assumption beneath much of enterprise AI marketing, which is that a product can speak the language of consequence before it has earned the right to carry one. The cleanest way to say it is also the one most likely to irritate the people selling the category. 2025 was the year of agent washing. 2026 is the opening act of outcome washing. By 2027, if the incentives now in view hold, nearly every large software company will be calling itself an outcome-based service, an outcome engine, a finished-work platform, or some other phrase designed to sound closer to consequence than software used to sound. Some of those claims will be earned. Many will not. A large share will be the same old products wrapped in a language model, workflow automation, and cleaner copy, then sold back to the market as if a change in wording had converted assistance into accountability. That is not a complaint about ambition. It is an observation about pressure. When the economics of a category start getting questioned, the story starts moving faster than the product.

The stronger sellers are not all the same

It would be a mistake to treat every company now speaking in the language of outcomes as if it were merely repainting weak software. Some vendors are making much stronger claims than simple copilots or retrieval layers. They describe systems operating inside live business workflows, executing transactions, coordinating multiple agents, automating operational steps, and acting directly in systems of record. In supply chain, finance, and operations, the claim is no longer only that software can suggest the next move. It is that software can create orders, trigger maintenance flows, update records, and carry multi-step work across a process that once depended on handoffs between people. Those are serious claims, and they should be treated seriously. They move well beyond the familiar pattern of a chatbot layered over documents or a copilot producing recommendations for human review. They suggest software that takes part in the work itself. Public descriptions from the stronger incumbents show direction, architectural ambition, workflow placement, orchestration, and the possibility of real intervention. One integrated suite vendor now says its agentic systems can make autonomous decisions about how to achieve a goal and execute those decisions, that its agents can take actions based on role and experience, and that they can be embedded in business processes and transactions where they automate standard steps and coordinate across workflows. That is not the language of chat, retrieval, or simple assistance. It is the language of action. But serious claims are not the same thing as demonstrated proof. What the public record shows more clearly is intent and positioning, not yet repeated evidence that these systems reliably change consequential business outcomes under accountability. There is a difference between participating in the workflow and carrying enough context, permission, and causal understanding to shape the result that matters. That difference is exactly where the market is now headed, and it is why the definition of an agent has to remain disciplined even when the product pages do not.

A stronger claim deserves a harder standard, not a softer one

Last year the market wanted a single word that could compress the promise of AI into something boards, analysts, and operating teams could recognize as new. “Agent” did that work. It made old software sound closer to action. It made copilots sound less advisory. It made automation sound less brittle. It made retrieval, orchestration, and routing feel like steps on the way to delegated judgment. The problem was never that the word had no use. The problem was that it got stretched until it stopped distinguishing one thing from another. A chatbot can retrieve. A copilot can assist. A workflow engine can route. An automation can execute steps. A multi-step orchestration layer can coordinate several of those motions at once. None of that, by itself, means the system is an agent in the sense that matters. If it cannot shape an outcome, it is not an agent. That line matters because it returns the burden of proof to the place the market worked so hard to avoid. An agent must be able to shape an outcome. Otherwise it is not an agent. That does not mean it must run wild, override human judgment, or operate outside governance. It means something simpler and harder. It must be able to act where the work happens, with enough context to understand the work, enough permission to intervene in the work, enough operational connection to do more than comment on it, and enough feedback from reality to know whether the intervention changed the result that matters. Without those conditions, the software may still be useful, profitable, and worth deploying. It is just not the thing the label implies. That was the weakness hiding inside much of the market’s language in 2025. It was easier to rename than to redesign. It was easier to add a model, a conversation layer, and a few orchestrated workflows than to move authority, rewrite permissions, harden feedback loops, and accept the legal, operational, and political burden that comes with software acting closer to consequence. So the category did what categories often do. It borrowed a harder word before it had earned the harder mechanism. The move made sense for a while. Markets reward novelty before they reward precision. Buyers wanted to believe the long wait between signal and action was finally getting shorter. Investors wanted to believe the old software stack could be revalued upward instead of repriced downward. Product teams wanted to believe a model atop a large installed base was enough to defend the moat. But once attention moved from demos to budgets and from budgets to results, “agent” stopped being a halo and started becoming a liability. The word invited questions too many products could not answer cleanly. What result did it actually change. What system could it write to. What did it decide without a human carrying the real burden. What evidence showed the effect persisted after the demo ended.

A safer word buys time

The safer word is “outcome.” It is easy to see why. Outcome is closer to what the buyer actually wants. A CFO does not want a better prompt. A COO does not want a cleaner chat pane. A board does not want a claim about digital labor unless it turns into service level, cycle time, uptime, yield, margin, cost-to-serve, working capital, or risk. Outcome language lets the seller move closer to the buyer’s real concern without having to defend every implication that agent language triggered. It lets the company say, in effect, do not focus on whether this system is truly autonomous. Focus on what it delivered. That is why the language coming out of the market matters. Microsoft is not merely telling executives to ask models better questions. It is telling them the next phase will be systems that plan, act, verify, revise, and deliver. Workday is reporting that employees are seeing time savings from AI while nearly 40 percent of those gains are being given back to rework, including correcting errors, rewriting content, and verifying outputs. A large enterprise suite vendor is now describing agentic outcomes and outcome-driven suites. ServiceNow speaks of mission-critical outcomes at scale. Salesforce is counting “agentic work units” and tying those counts to revenue momentum. The names differ. The direction does not. The category is turning from intelligence as spectacle toward consequence as marketing claim. None of this is meaningless. Some of these products are getting better in ways that matter. Some can already execute bounded actions with real effect. Some are more deeply embedded in systems of record than the market gives them credit for. Some really are carrying more burden than the software they replace. That counterexample has to be treated fairly or the argument becomes too easy. The integrated stack players have more room than thin wrappers do to ground AI in shared data and permissions. Salesforce’s numbers at least show there is market demand for systems sold as doing more than talking. Microsoft is right that the next contest is not model quality alone but the design and governance of systems that can execute end-to-end work. The point is not that nothing has improved. The point is that the market has a habit of claiming the result before the burden has moved far enough to justify the claim.

The buyer pays twice when the burden stays put

For most of the SaaS era, the bargain was plain. The vendor sold access to capability. The customer bought the platform, connected it to the rest of the stack, trained people to use it, bent process around it, and then spent labor to turn access into outcome. The software was valuable, often enormously so, but the customer still carried much of the consequence. A report did not create the result. A ticketing system did not resolve the customer issue by itself. A dashboard did not restore uptime. An enterprise suite did not remove the need for managers, coordinators, reviewers, escalations, and handoffs. The software made those efforts more orderly and more visible. It did not carry them for you. AI has made that bargain look less complete. Once software can read, draft, classify, retrieve, route, and in some cases act, the buyer starts asking a tougher question. If the product is now closer to the work, why am I still carrying so much of the labor between knowing and doing. Why do I still need so many intermediaries between a signal and a correction. Why am I still

paying the meeting tax, the handoff tax, the review tax, the escalation tax, and the coordination tax while being told that the system is now intelligent. That is the pressure beneath the rhetoric in early 2026. It is not only pressure on valuation. It is pressure on the legitimacy of access-based software economics in a market that is starting to expect burden transfer. That is why the selloff in software stocks in February mattered beyond the market-cap number itself. Reuters reported that nearly $1 trillion in value had been wiped from software and services stocks over a week as investors debated whether AI posed an existential threat to the business models of traditional software companies. The phrase can be overused. Here it was not entirely wrong. The threat is not that software disappears. The threat is that customers begin to pay differently once they believe the old bargain no longer matches the new possibility. If the product claims to carry more of the work, the buyer will increasingly ask why it is still being paid as if the customer must carry nearly all of it. If the product does not carry more of the work, the buyer will eventually ask why it is being described in language that implies it does. Either way, the pressure lands on the same fault line. Who is still carrying the burden. The buyer no longer wants software that explains the work. The buyer wants software that carries more of it. That is also why the market’s wording is getting more practiced. It is not becoming more honest. It is becoming more skilled at moving one step ahead of the question it cannot answer cleanly. Last year that meant renaming assistive systems as agents. This year it increasingly means leading with delivered work, measurable results, and outcomes at scale. The mechanism underneath may be improving. The language above it is almost certainly improving faster.

The test is in the burden, not the copy

The problem for buyers is that this turn from agent language to outcome language does not make the market easier to read. It makes it harder. Once every vendor begins claiming results, the only question that matters is whether the software actually carries the burden of producing the outcome or whether that burden still sits with the human organization. A genuine outcomeshaping agent must be able to change a business result without a human deciding first. It must be able to act inside the real systems where the work happens, not simply analyze data, generate text, or recommend the next step. That means writing into systems of record and executing actions that alter the state of the workflow itself. It must also operate with delegated decision rights. If every consequential step still requires human approval, then the system is not shaping the outcome. It is assisting, routing, or accelerating human work. Just as important, it must understand what intervention is likely to change the KPI that matters and why. Predicting patterns is not enough. Finally, it must operate in a closed loop. It must observe the signal, choose an intervention, execute the action, measure the result, and learn from what actually happened. If a system cannot intervene in the workflow, cannot execute actions in the operational systems where the work lives, cannot explain why its action should change the result, and cannot verify whether the result actually improved, then the

burden of shaping the outcome still belongs to the enterprise. In that case the software may still be useful, but it is not an agent shaping outcomes. It is still a tool helping people carry the work. This is where the article should stop being theoretical for the reader. If a product cannot write to the system of record where the relevant work lives, is it really shaping the outcome or merely commenting on it? If it cannot act without a human deciding first, what exactly has been delegated besides language? If it cannot explain, in plain operational terms, why this intervention should change this KPI under these conditions, what is the buyer trusting besides fluency? And if it cannot verify what happened after the action, then learn from the consequence, what part of the claim survives contact with the work? A second paragraph deserves to be read aloud in any executive meeting where someone says the result has already arrived. What business result can this product change without a human making the decision first? Which system does it actually act inside? What part of the burden moves from payroll to platform, and what part merely gets described more elegantly in the demo? If the answer remains cloudy, the claim remains early. If the hidden human work would rush back tomorrow the moment the product disappeared, then the software may have reduced friction but it has not absorbed enough consequence to justify the larger story being told about it.

Outcome washing starts when the claim outruns the burden transfer

Agent washing was easy to spot once the novelty wore off. The product page would use hard words. The workflow would remain soft. The system would appear active but depend on quiet, constant human correction. The “agent” would route, draft, and summarize, while the enterprise still made the consequential decisions and absorbed the downside. Outcome washing is subtler because it does not fight the buyer’s motive. It flatters it. It says, correctly, that the buyer does not care about the model for its own sake. It says, correctly, that what matters is the result. Then it uses that truth to lower scrutiny of what the product actually had to do to deserve the word. This is how it works in practice. The vendor stops leading with what the system is and starts leading with what it says the system delivers. The product is described less as an agent and more as a work-completion engine, a finished-results platform, or an outcome layer across the enterprise. Time saved becomes value created. Tasks completed become burden removed. Work units become evidence of consequence. The seller speaks less about autonomy because autonomy invites inspection. It speaks more about results because results, if left unexamined, let several forms of assistance hide under one polished term. That is the opening through which old products can be sold as new categories. The buyer’s risk is not just overpaying for hype. The deeper risk is losing the language needed to audit what the system is really doing. Once outcome becomes a catch-all term, anything that contributes to a result can be described as producing one. A draft that a human must rewrite becomes part of the outcome. A recommendation that still waits for a manager’s approval becomes part of the outcome. A workflow that surfaces the next best step while an operator still executes the real step becomes part of the outcome. The software begins collecting narrative

credit for work the enterprise still carries in payroll, in meetings, in management load, and in cleanup. Workday’s January research is useful here because it exposes the hidden mechanism so cleanly. Nearly 40 percent of AI time savings were being lost to rework. That means the market already has evidence that fast output is not the same thing as a net result. The buyer who saves ten hours and gives back four in correction has not found a clean outcome. The buyer has found a gross speed gain with a concealed tax on verification. That is not failure. It is a reminder that the product’s work and the enterprise’s work are still entangled. The burden did not move far enough to justify triumph. Salesforce’s phrase “agentic work units” is useful for a different reason. It is a serious commercial attempt to count digital labor. That may prove valuable. It may also teach the market to confuse the vendor’s internal unit of activity with the buyer’s business result. A work unit is not automatically margin. It is not automatically retention. It is not automatically yield, uptime, or faster cash conversion. A software company will always have a reason to count what flatters the architecture it has built. The buyer’s job is to keep asking whether the unit being counted is the same as the result that matters. The market will keep polishing the label. The burden will tell the truth. There is a fair counterargument and it has to be addressed plainly. Some of what is now being sold may become true in bounded domains sooner than critics expect. The installed-base vendors do have real advantages in data access, workflow depth, identity, and permissions. In narrow domains, with limited authority boundaries and clean feedback, some of these products may deserve stronger words than the market had available a year ago. That possibility is real. But it does not rescue the category from the discipline it is trying to avoid. It sharpens the need for it. If some products are moving into outcome-shaping territory, buyers need harder standards, not softer ones. They need to tell the difference between real burden transfer and cleaner interface design. They need to know where delegated action is real, where it is partial, and where it is still theatre wrapped around human review. A stronger market needs more precise language, not more forgiving language. There is a second counterargument worth treating honestly as well. Sometimes the vendor is not the only one avoiding consequence. Enterprises themselves often want the language of outcome without paying the organizational price required to let software carry one. They keep permissions fragmented, decision rights vague, systems disconnected, and accountability so politically charged that no product could be trusted to act cleanly even if it were technically able to do so. In those cases the software may be ahead of the operating model, not ahead of reality entirely. That is true. It is also beside the point as the claim is sold. If a product can only shape outcomes after the buyer redesigns governance, workflow ownership, authority, and consequence management, then that should be named plainly. It should not be sold as if the result is already in hand.

The claim will spread faster than the proof

Here is the prediction that would be embarrassing if wrong. By the end of 2027, the dominant language on major enterprise AI product pages, earnings calls, and executive keynotes will not be agents alone. It will be outcomes, delivered work, finished results, mission-critical results, or some close variant. If that does not happen, this argument deserves to be treated as overdrawn. But the evidence already points the other way. The rhetoric is tightening around results because results are where buyer scrutiny is headed. The category is learning to speak in the grammar of burden transfer, whether the burden has moved enough or not. That matters because a market can survive inflated language longer than buyers can survive surrendering the words needed to resist it. The danger is not merely paying too much for software. The danger is paying for a narrative that says the enterprise is finally free of old frictions while the real costs continue to arrive, only now in harder-to-see forms. Rework. Verification. Human override. Approval bottlenecks. Exception handling. Escalation. Quiet management labor. Those costs do not disappear because a product page says the outcome has been delivered. They disappear only when the burden itself moves. A pile of documents with an LLM on top is still a pile of documents with an LLM on top. That sentence will sound too harsh to some people. It is not harsh enough. The point is not that document-grounded systems are useless. The point is that the market has repeatedly tried to cross a semantic bridge that the mechanism beneath it had not yet built. An LLM over content can be very helpful. Add workflow automation and it can be more helpful. Add routing, retrieval, and orchestration and it can be more helpful still. None of that is the same as outcome-shaping capability unless the system can intervene where the result is formed, with permission to do so, and with feedback strong enough to learn from the consequence. Without that, the architecture remains closer to a recommendation engine wrapped in enterprise ceremony than to a true carrier of outcome burden. Three months into 2026, this is where the market stands. It is not beyond hype. It is in a more disciplined phase of it. The words are getting better because buyers are getting harder to fool. Sellers know they cannot rely on “agent” alone forever. So the new vocabulary moves one step closer to the boardroom. Outcome. Result. Work delivered. Mission-critical consequence. Completed business action. The market will get more polished from here. That is why the buyer must get less impressionable. There is only one durable defense. Keep the definition fixed. An agent must be able to shape an outcome. Otherwise it is not an agent. Then ask the question that follows from it every time a vendor says the result has arrived. Who is still carrying the burden. If the answer is still mostly the enterprise, then the category has changed its wording faster than it has changed its truth. The label will keep moving. The burden will not. References

This piece draws on the current public record because the argument is not only conceptual. It is already in the market. Reuters reported on February 4 and February 5, 2026 that software and data-services stocks had lost roughly $1 trillion in value as investors weighed whether fastmoving AI tools threatened the economics of the sector. Microsoft WorkLab’s February 17, 2026 article matters because it states plainly that the next phase of AI at work is systems that plan, act, verify, revise, and deliver, which validates that the market is moving from output toward execution. Workday’s January 14, 2026 research matters because it found nearly 40 percent of AI time savings are lost to rework, which is direct ballast for the claim that gross speed is not the same as clean result. Salesforce’s February 25, 2026 fiscal 2026 results matter because they quantify Agentforce at $800 million in ARR, 29,000 deals, and more than 2.4 billion agentic work units, which shows both real demand and the temptation to let vendordefined units stand in for business consequence. Oracle’s public material on agentic AI, AI agents, AI for supply chain and manufacturing, and AI Database 26ai matters because it shows that at least some incumbents are making stronger claims about action, orchestration, transaction execution, and workflow placement than simple wrapper software ever could, even if public proof still trails those claims. As conceptual ballast beneath the current record, Deming’s Out of the Crisis in 1986 remains essential on quality and the cost of rework, Herbert Simon’s work on bounded rationality still explains why overloaded decision systems hide latency until it becomes structural, Judea Pearl’s The Book of Why in 2018 still matters because intervention is not the same as correlation, and Ronald Coase’s 1937 theory of the firm still matters because coordination cost, not software rhetoric, decides where burden actually sits.

Topics: synthetic-agency, causal-aiOpen in the Radiant ↗All dispatches