Everyone Is Trying Agents; Almost No One Has Scaled Them
The headline numbers on AI adoption have stopped being interesting because they are all near-universal. According to McKinsey’s 2026 State of AI global survey, 88% of respondents say their organizations regularly use AI in at least one business function, and 72% report using generative AI — up sharply from 33% in 2024. Adoption, in the broad sense, is settled.
The interesting number is the one underneath. McKinsey finds that 62% of organizations are at least experimenting with AI agents, and 23% report scaling an agentic AI system somewhere in the enterprise — but in any single business function, as CX Today’s analysis of the report notes, no more than 10% of organizations are scaling agents. The distance between “experimenting” and “scaled in production” is the defining enterprise AI problem of the year, and it is wide.
What “Scaling” Actually Requires
The reason the gap is so persistent is that scaling an agent is a fundamentally different exercise from piloting one. A pilot succeeds when a single agent completes a single workflow in a controlled setting for a small group of users. Scaling requires that the same agent run reliably across many workflows, for many users, against live production data, under governance, with observable behavior and predictable cost. Each of those requirements is a separate engineering and organizational commitment.
McKinsey’s data makes the payoff conditional on exactly this kind of rework. The survey finds that the small group of organizations drawing meaningful, scaled value — roughly 6% of respondents qualify as “AI high performers” who attribute more than 5% of EBIT to AI — behave differently from everyone else. As CX Today reports from the survey, 55% of high performers say they have fundamentally redesigned workflows to capture AI value, versus only 20% of other organizations, and 65% have defined human-in-the-loop validation processes, versus 23%. The advantage is not the model; it is the operating discipline around the model.
The EBIT Reality Check
The most sobering figure in the survey is the value one. Despite 88% regular usage, McKinsey’s survey finds only 39% of organizations report any EBIT impact attributable to AI — and only about 6% draw a meaningful, scaled EBIT impact. That is a large majority of companies spending on AI without yet being able to point to it on the income statement.
This is not evidence that AI does not work; the high performers demonstrate it does. It is evidence that value accrues to the organizations that treat AI as an operating-model change rather than a feature to bolt on. The 39%-with-any-impact figure and the ~6%-with-scaled-impact figure describe the same reality from two angles: most organizations have adopted the tools but not the discipline that turns tools into margin.
Advertisement
Why Agents Are Harder Than Generative AI Was
Generative AI diffused through enterprises quickly because it slots into existing human workflows — a person prompts a model, reads the output, and decides what to do with it. Agents remove the human from the middle of the loop, and that is exactly where the difficulty concentrates:
- Autonomy raises the stakes of every error. A chatbot that hallucinates wastes a user’s time. An agent that hallucinates while executing a multi-step workflow can take real, wrong actions against production systems before anyone notices.
- Evaluation is unsolved. You can grade a summary. Grading whether an autonomous agent made a good sequence of decisions across a long task is far harder, and the absence of trusted evaluation is a primary reason pilots stall.
- Governance friction is real. Deciding which actions an agent may take, on whose behalf, with what data, and with what human sign-off, is organizational work that touches security, legal and compliance — not something an engineering team ships alone.
- Cost is non-linear. A single agent that spawns sub-agents and retries can generate orders of magnitude more model calls than a pilot suggested, and traditional FinOps dashboards report the bill after the fact.
None of these are reasons not to deploy agents. They are the specific reasons the 62%-experimenting cohort has not become a 62%-scaled cohort — and the checklist any organization must work through to cross the gap.
What This Means for Enterprise Leaders
The survey’s implicit prescription is that scaling is an operating-model project, not a procurement decision. Leaders serious about moving from pilot to production should make the following structural moves.
1. Redesign the workflow around the agent before deploying it
High performers redesign workflows; laggards bolt agents onto unchanged processes. Before deploying an agent, map the end-to-end workflow it will run, decide which steps it owns and which stay human, and rebuild the process around that division — measuring against a real baseline so you can prove impact later.
2. Build agent evaluation before you build the agent
The reason pilots stall is that no one can confidently say the agent is good enough. Define, up front, how you will measure success across full task sequences — not just single outputs — and stand up the evaluation harness before scaling. If you cannot measure it, you cannot responsibly scale it.
3. Install human-in-the-loop validation as a control, not an afterthought
65% of high performers have defined human-in-the-loop validation; most others have not. Decide which agent actions require human sign-off, build those checkpoints into the workflow, and assign a named owner accountable for each agent’s behavior in production — before deployment, not after an incident.
4. Instrument cost and governance from day one
Route agent actions through a policy layer that enforces permitted tools, data-access rules and rate limits, and put per-agent cost and behavior into an observability system your security, finance and engineering teams can all query. Non-linear cost and autonomous action are only manageable if they are visible.
The Correction Scenario
The scaling gap is unlikely to close by everyone suddenly becoming a high performer. The more probable 2026-2027 path is bifurcation: the ~6% who have rebuilt their operating model pull further ahead, capturing disproportionate EBIT, while the long tail continues to run pilots that never graduate. For that long tail, the risk is not falling behind on technology — the models are broadly available — but falling behind on the organizational discipline that turns the models into results.
For enterprises in emerging markets, including Algeria, this framing is unexpectedly encouraging. The gating factor McKinsey identifies is not access to frontier models or hyperscale budgets; it is workflow redesign, evaluation and governance — capabilities that a disciplined mid-sized organization can build without a nine-figure AI program. The lesson of the 2026 survey is that the winners are separated from the rest less by what they bought than by how they rebuilt the work around it.
Frequently Asked Questions
What is the AI agent scaling gap?
It is the distance between organizations that experiment with AI agents and the far smaller number that run them in production at scale. McKinsey’s 2026 State of AI survey found that 62% of organizations are at least experimenting with AI agents and 23% are scaling an agentic system somewhere, but in any single business function no more than 10% of organizations are actually scaling agents.
Why do so few enterprises capture value from AI?
According to McKinsey’s 2026 survey, only about 39% of organizations report any EBIT impact from AI, and only roughly 6% draw a meaningful, scaled EBIT impact. The gap is driven less by the technology than by operating discipline: high performers redesign workflows (55% versus 20% of others) and define human-in-the-loop validation (65% versus 23%), while most organizations bolt AI onto unchanged processes.
How can an organization move an agent from pilot to production?
The survey points to four moves: redesign the workflow around the agent rather than bolting it on, build an evaluation harness that measures full task sequences before scaling, install human-in-the-loop validation as a control with a named owner, and instrument cost and governance from day one. McKinsey’s data shows these operating-model changes — not access to better models — separate high performers from the rest.














