
Deloitte's 2026 State of AI in the Enterprise survey found that 66% of companies report real productivity and efficiency gains from AI. Only 20% report revenue growth from it — even though 74% say revenue growth was the goal. That's a 46-point gap between what AI delivers and what it was funded to deliver, and it's well enough documented now to ask precisely why it exists, not just that it does.
Three of the year's most-cited enterprise AI studies converged on the same pattern independently. McKinsey's State of AI 2026 survey (1,719 executives, 97 countries, fielded May–June) found that while roughly nine in ten companies use AI somewhere in the business, only 37% can attribute any EBIT impact to it, and just 6% qualify as "AI high performers" — a 5%+ EBIT contribution. Bain's parallel survey found half of companies realizing only 0–10% in cost savings, roughly half of what they projected, and yet 90% are increasing AI budgets anyway. Deloitte's efficiency-versus-revenue split adds the sharpest detail: the gains sit almost entirely on the cost side of the P&L.
This built on MIT NANDA's August 2025 "GenAI Divide" report, whose finding that 95% of custom enterprise GenAI pilots showed zero measurable ROI against an estimated $30–40 billion in investment became the reference point nearly every subsequent 2026 story cites.
For engineering leaders, this isn't a marketing problem — it's an architecture and prioritization problem with a direct line to what gets built, staffed, and funded next quarter. Boards approved AI budgets in 2024–2025 on the promise of growth. What's landing on P&Ls in 2026 is cost avoidance in bounded workflows, and CFOs now underwrite the next AI proposal like any capital project — against a specific opex, margin, or earnings hypothesis, not a productivity anecdote. Teams that can't show which P&L line a project affects, and by how much, get cut first when this year's numbers come up for review.
The mechanism is simpler than it looks, and Bain names it directly: most organizations are "layering" AI onto an existing process rather than redesigning around it. Layering caps the outcome at "do the current thing more cheaply," which shows up cleanly as a cost line, because the baseline (headcount hours, agency spend, ticket-handling time) was already being measured. Revenue growth requires a different kind of change — a new product, segment, or pricing model — something with no pre-existing baseline to improve against, and that requires product decisions AI alone doesn't make.
MIT NANDA's data makes the same point from the deployment angle: external partnerships with real customization succeeded 67% of the time versus 33% for internal builds, and the highest-ROI cases were narrow back-office automations (BPO elimination, agency-spend cuts, risk-check automation) rather than broad front-office "transformation" bets. Narrow and well-scoped beats broad and aspirational — and narrow scope is exactly what produces a cost number instead of a revenue story.
Demonstrated, evidence-backed: Individual productivity gains are real and broad — 80% of AI users report improved personal output (McKinsey). Bounded automation produces measurable savings: BPO elimination worth $2–10 million a year, 30% cuts in external agency spend, and roughly $1 million a year in financial-services risk-check savings (MIT NANDA), achieved largely without material workforce reduction.
Emerging, not yet conclusive: Large enterprises (above $1 billion revenue) are scaling AI faster than mid-market firms — 54% report enterprise-wide scaling versus roughly a third of smaller companies — but it isn't clear yet whether the return gap closes faster for large firms or simply gets funded longer while it stays open.
Open debate: McKinsey's own research team frames the gap as timing, not a value ceiling — an "AI productivity J-curve," drawing the historical parallel to factory electrification, where intangible investment (process redesign, retraining) preceded any output gain by years. Bain's "layering" diagnosis is consistent with that data but implies a different prescription: most AI investment needs re-architecting now, not more time. Neither source rebuts the other, and patience has a cost either way — Gartner projects at least half of GenAI projects will overrun budget through 2028 on poor architecture, and agentic workflows already consume 5–30 times the inference tokens of a standard chatbot.
The Klarna case is worth naming because it's concrete rather than statistical. The company's 2024 claim of replacing roughly 700 support staff with AI for an estimated $40 million a year in savings didn't hold up once complex billing disputes, fraud cases, and emotionally charged interactions drove enough repeat contacts to erode the savings. Klarna quietly moved to a tiered human-plus-AI model through 2025–2026 — a cost-automation architecture applied to judgment-heavy work that needed a different one.
The pattern is consistent enough to draw a practical conclusion: the ROI gap is not primarily a model-quality problem. It's an implementation-discipline problem, and that's a solvable one.
The organizations MIT NANDA classifies as high performers share two habits that show up repeatedly across the other studies: they redesign the workflow before automating it, and they treat external delivery partners as accountable for the operational outcome, not just for shipping an integration. Both are architecture and vendor-management decisions an engineering organization controls directly — neither requires a better model to arrive.
Agentic AI deserves more scrutiny than the current enthusiasm suggests. It's the newest, least governed, and most token-intensive category here, and Deloitte found only one in five companies has a mature governance model for autonomous agents. Given the token-consumption multiplier Gartner and McKinsey both cite, cost governance reads like a decision worth making at design time, before the usage pattern is locked in — not something to work out once a rollout is already generating a monthly bill.
The J-curve argument is plausible but unproven for any specific deployment — a reasonable case for patience with a well-architected initiative genuinely early in a redesign, not a justification for continuing to fund a bolted-on automation that was never going to touch revenue in the first place.
Architecture. Scope discipline correlates with realized ROI. A narrow workflow with a real cost baseline (support triage, document processing, risk checks) is measurable and fundable. A broad "transformation" initiative without a specific process being redesigned tends to land in the 90% of use cases McKinsey found stuck in pilot.
Cost. Roughly one in five organizations already say AI operating costs constrain usage. For agentic deployments, model the token-consumption profile against expected volume before committing to production — the multiplier over standard chat is large enough to change project economics.
Governance. Autonomous deployments need an owner, an audit trail, and a defined kill criterion before they go live, not after an incident — the gap Deloitte's governance-maturity finding points to directly.
Observability. Attribute financial impact at the workflow level, against a documented pre-AI baseline, before scaling. No survey here directly tested what separates the 6% of "high performers" from everyone else, but baseline measurement shows up consistently across the organizations MIT NANDA and McKinsey both credit with real EBIT impact.
This isn't a single yes-or-no decision — the evidence supports different postures by use case.
Adopt narrow, back-office automation with a documented before/after cost baseline — the strongest, most consistent evidence here.
Experiment, with explicit cost governance and a kill criterion, for agentic and revenue-facing initiatives — the weakest evidence, least predictable economics, and lowest governance maturity industry-wide.
Monitor the J-curve thesis over the next few quarters before treating "revenue transformation is coming" as a safe planning assumption. What would change this: a second consecutive survey cycle showing the EBIT-impact number actually move, rather than holding flat as it did between 2025 and 2026.
The gap between AI efficiency and AI revenue isn't evidence the technology doesn't work — the productivity numbers are real. It's evidence that most organizations are measuring a redesigned-workflow outcome against an unredesigned-workflow architecture, and getting exactly the cost-line result that architecture was always going to produce. The 6% getting a different result changed the workflow first and the tooling second.
If your organization can point to AI-driven productivity gains but can't yet say what they've done to margin, that's usually a visibility problem before it's a technology problem. Lestar AI CFO Assistant gives finance leaders a direct read on growth, margins, and liquidity — built for teams still waiting on Finance-produced spreadsheets to answer whether an investment, AI or otherwise, is showing up in the numbers. We typically structure engagements around measurable checkpoints, with initial results visible within 6–12 weeks and full ROI assessed within 6–12 months of deployment. Talk to Lestar.
Five independently sponsored 2026 surveys agree on an uncomfortable number: only 5 to 7 percent of enterprises say their data is genuinely ready to support AI at scale, even as nearly all of them report active AI initiatives. The gap gets sharper with agentic AI, where ungoverned data doesn't just produce a wrong number on a dashboard — it produces a wrong action. This piece looks at what a testable definition of "AI-ready" actually requires, and why self-assessment keeps missing the difference.
Three-quarters of enterprise leaders say they've adopted agentic AI; independent research puts verified production ROI closer to five percent. This piece unpacks why those numbers aren't actually in conflict, and what separates the deployments compounding at scale from the ones Gartner expects will be canceled by 2027.
Malaysia's AI Governance Bill hasn't passed Parliament, and its penalty schedule hasn't even been written — but the shape of its obligations already has. This piece breaks down the Developer/Deployer split, the overlap with existing PDPA law, and why engineering teams building or deploying AI in Malaysia have reason to start preparing before the law is finalized.
Whether you need Lestar ESG for sustainability reporting, Lestar CEO360 for executive intelligence, or a fully customised enterprise data implementation — Mandrill Tech will tailor the solution to your organisation's needs.