On August 26, 2025, MIT's Project NANDA published a report titled "The GenAI Divide: State of AI in Business." One line drove the headline for the next twelve months. Ninety-five percent of enterprise AI pilots delivered no measurable P&L impact.
A year on, the number is still cited. Fortune, Forbes, HBR, every consultancy deck. What almost nobody has done in that year is change the metric they use to justify the next pilot.
The GenAI Divide finding rested on interviews with 52 executives, surveys of 153 leaders, and analysis of 300 public AI deployments. In 2026, the follow-on data has only sharpened it. S&P Global Market Intelligence, in a survey of more than 1,000 enterprises across North America and Europe, reports that 42% of companies abandoned most of their AI initiatives in 2025, up from 17% the year before. The average organization scrapped 46% of its AI proof-of-concepts before they reached production. Gartner puts the share of AI agent pilots that never ship to production at 88%.
Meanwhile, WRITER's 2026 Enterprise AI Adoption survey reports that 91% of businesses use AI and 72% have at least one AI deployment live. Anthropic's Economic Index, which now publishes actual usage data drawn from Claude, shows Claude Code alone crossed two million weekly active users this year, with clear working-hours seasonality across white-collar knowledge work.
Two curves. One climbs. One goes flat. And the flat one is the one your board actually reads.
The chart went up. The money did not.
Every enterprise AI dashboard I have seen in the last twelve months tracks the same four numbers. Seats deployed. Weekly active users. Prompt volume. Percent of employees who have logged in at least once. These are activity metrics. They measure whether people touched the tool. They say nothing about whether the tool did any work.
The Forbes analysis of CEO returns published in January 2026 found that 56% of chief executives saw neither increased revenue nor decreased costs from AI in the trailing twelve months. Only 12% reported both. Twenty-nine percent of executives said they can confidently measure AI ROI at all. That gap is the whole story. Adoption climbed. Return did not. The chart kept going up because the chart was measuring something adjacent to value, and not value itself.
There is a name for this shape in operations. It is a vanity denominator. You pick a number that is easy to move, you push on it, and you point at the movement to justify the budget. The 91% adoption number is a vanity denominator. So is the two-million-weekly-active number. So is your internal Copilot deployment count. Every one of those can climb every quarter while the P&L line for AI stays flat, and the person defending the budget will still get a hearing.
What is not a vanity denominator is the outcome the pilot promised in its charter. The MIT number tells you that ninety-five out of every hundred pilots produced none of that outcome.
The AI adoption metrics that matter
If you are running the AI portfolio for a company that has a board, and that board wants to see the return, the AI adoption metrics that matter are a different set of numbers than the ones your CIO is reporting today. Five of them carry weight.
Evaluation coverage per change
Among enterprises that shipped AI to production and kept it there, 87% run automated evaluations on every prompt change, every model swap, every tool update. Not spot checks. Every change. That coverage number tracks discipline. Companies that ship and companies that do not are separated less by which model they picked and more by whether they built the eval harness that catches regressions before the customer does.
Seventy percent of leaders name "non-deterministic outputs" as the top production-readiness barrier. That is a polite way of saying they cannot tell when the model got it wrong until a user complains. Evaluation coverage is the metric that turns that fog into something a change-management process can operate on. Without it, every model provider update becomes a silent liability. With it, every model provider update becomes a routine test run.
Ask your team what percent of production prompts have automated evals with a passing baseline. If the answer is under 50%, you do not have an AI portfolio. You have a set of experiments that happen to be running in production.
Time-to-measurable-outcome
The median AI initiative that eventually shows P&L takes 14 months to get there. Fourteen. That number kills the CFO's quarterly-review reflex to defund a pilot at month four when it has not hit the number yet.
It also disciplines the ask. If your team cannot articulate what the measurable outcome is on day zero, they will not find it on day 400. A pilot that starts with "we will cut ticket-handle cost by $2M in twelve months, measured as ticket-handle time times fully loaded labor rate" is a pilot with a clock and a finish line. A pilot that starts with "we will explore generative AI for support" is a pilot that will appear in the S&P Global abandonment column next year.
Track the age of every live pilot and the specific dollar hypothesis attached to it. Any pilot older than three months without a written outcome hypothesis gets closed or rewritten. That is a metric a board can act on inside one meeting.
Cost per resolved outcome
Every AI feature you run has a unit economic. Cost per resolved support ticket. Cost per underwritten loan. Cost per drafted contract. Cost per completed sales-development cycle. The numerator is inference cost plus human-in-the-loop labor. The denominator is the count of outcomes actually resolved end to end, with no human rework beyond a defined threshold.
Almost no one tracks this. Internal reporting stops at "we spent $340k on Claude API this month" and never divides that by the count of things the system finished. When you do the division, you learn which features are cheaper than the human process they replaced and which ones are more expensive than the human process they replaced.
The public benchmarks are already sharp. Fin AI's 2026 industry data puts AI-resolved customer support tickets at $0.99 to $2.00 per interaction against $6 to $12 for human handling. Reddit's Salesforce Agentforce 360 deployment deflected 46% of support cases and cut resolution time by 84%. That is a real gap. It only matters if you know what your specific per-resolved-outcome number is, versus your specific human-baseline number, and can defend the difference to the CFO.
The MIT number reads differently under this lens. Ninety-five percent of pilots showed no P&L impact because ninety-five percent of pilots were never running the unit economic math at all. They were running usage math.
Workflow depth
A tool bolted onto the side of a workflow gives you a seat count. A tool inside the workflow gives you a P&L line. This distinction lives in measurement.
The MIT NANDA report specifically found that the 5% of pilots that delivered value shared a pattern. They designed for friction, embedded into a high-value workflow, integrated deeply, and shipped tools that carried memory across sessions. The 95% that failed treated GenAI as a productivity garnish sprinkled on top of existing knowledge work. The garnish had adoption. It had prompt volume. It had a weekly active user chart. It had a workflow that never changed.
The measure here is a ratio. What percent of the completion of a bounded, valued workflow now happens inside the AI system? If your customer-service automation covers 4% of ticket completion end to end, the tool is cosmetic. If it covers 55%, the tool is doing the work. Workflow depth is the single ratio that predicts whether the CFO will keep funding the line.
Killed-pilot velocity
This is the metric the industry does not want to hear. The 42% of enterprises that abandoned most of their AI in 2025 are being written up as failures. Some of them are the companies that will still be in business in five years.
An AI portfolio without a high killed-pilot velocity is an AI portfolio that is not learning. If you have twelve running pilots and none of them have died in the last two quarters, you are not running an AI program. You are running a museum. The organizations that make it across the GenAI Divide have a written process for killing a pilot at a defined checkpoint, and they use it.
The healthy version of the S&P Global abandonment number is a company that started fifty proof-of-concepts, killed forty-six of them on schedule at pre-declared decision gates, and moved the freed budget into the four that hit their unit economics. The unhealthy version is a company that started fifty proof-of-concepts, killed forty-six of them at random after they had already gone six months over budget, and cannot say why the survivors survived.
Track the pilot mortality rate and record a reason code on every kill. The reasons will show you which sectors of your business have real AI leverage and which ones do not.
What Klarna proves about the wrong single number
Klarna is the case study everyone quotes when they want a happy AI adoption story. Its OpenAI-powered support assistant handles roughly two-thirds of chats, equivalent to about 700 full-time agents. That was the 2024 press release.
The 2026 follow-up is different. Klarna publicly walked back an agent-only support deployment after admitting that cost had been given too much weight in the evaluation and quality had dropped. The company is hiring humans back into a hybrid model. This did not happen because the AI failed to work. It happened because Klarna had optimized to a single number, cost per contact, and shipped that number into production before a second number, customer-outcome quality on the harder half of the queue, had been given equal weight.
That is the whole essay in one company. Gartner's own data now finds AI deflects 45% or more of queries while only 14% of issues reach full self-service resolution. Deflection is a vanity denominator. Resolution is not. The gap between them is where the customer stops trusting the channel and the metric that got tracked stops predicting the metric that mattered.
Why the wrong metric got adopted first
Activity metrics are legible. You can put "12,000 employees using Copilot" on a slide and the board understands what it means. You cannot put "average cost per resolved incident dropped from $18 to $4 in six months, weighted by evaluation-pass rate, with 55% workflow depth" on a slide without losing three board members on the first phrase.
So the metric that got tracked was the one that fit on the slide. The metric that predicted survival was the one that did not.
There is a second reason. Seat count is a metric the vendor happily reports. Cost-per-outcome is a metric that would make the vendor look bad, because it forces the buyer to notice that inference is a variable cost and headcount reduction is a fixed saving and the two do not always net out in the buyer's favor. Your Copilot invoice grows with usage. Your labor savings do not grow with usage. Somewhere in the middle sits a crossover point that vendor dashboards are not built to show you.
The Anthropic Economic Index is a partial exception because Anthropic publishes actual usage data. The June 2026 report lets you see that a large share of Claude usage is code and text generation for individual white-collar knowledge work, with weekly seasonality that tracks working hours and drops on weekends. That is honest and useful. It is also an aggregate. It cannot tell you whether your specific deployment is producing P&L. It can only tell you that many people are typing at Claude.
The architectural read
The MIT report frames this as a divide between the 5% who cross and the 95% who do not. A framing that matters more to a CEO is a divide between companies who bought AI tools and companies who architected AI systems.
Buying a tool is a purchase decision. The metric that follows is a usage metric. Did people log in. Did they use it. How much did we spend on seats this year.
Architecting a system is a design decision. The metric that follows is an outcome metric. What did the system produce. What did it cost per unit of production. What percent of the target workflow does it now handle. What was the evaluation-pass rate on last week's changes. When we decide to kill it, at what threshold do we pull the plug.
Tools are billed per seat because the vendor wants to price the way that scales revenue with headcount. Systems are billed per outcome because the buyer wants to price the way that scales cost with value. These two purchases require completely different measurement discipline, and the org that bought the first while measuring like it bought the second is the org whose CEO just filled out the survey saying they got no ROI.
What to do this quarter
Take one pilot that is running now. Any one. Write four numbers next to it.
The evaluation-pass rate on its production prompts, measured against a baseline test set. The dollar cost per completed outcome, computed from your actual inference bill divided by your actual completed-outcome count. The percent of a named, bounded workflow that the system now handles end to end. The date at which you have committed to either scale it or kill it, and the specific threshold that decides which.
If you cannot fill in any one of the four, you have found the reason it is not producing P&L. The model is fine. The counting was where you lost the money.
The 91% adoption number belongs on a marketing chart. The four numbers above belong on a P&L. The distance between them is the GenAI Divide.
Architect the measurement into the deployment
The organizations that will survive this cycle share one habit. They keep a tight measurement loop between an outcome hypothesis, an evaluation harness, a unit-economic denominator, and a kill switch. That loop is not a feature you can buy from a vendor. It is an architectural decision you make before the first pilot ships.
That is what an AI advisory partner does. It sits with you before you sign the SaaS contract, defines the outcome hypothesis in language the CFO will still accept in month fourteen, builds the evaluation harness that lets you notice regressions before your customers do, sets the unit-economic denominator so you know whether the feature is cheaper than the human process it replaced, and installs the kill gate so the pilot that is not converging does not eat your Q3 budget silently.
The GenAI Divide is a measurement failure. Close it at the design of the next pilot, before the seat count starts climbing and the outcome number sits at zero.
Sources
- The GenAI Divide: MIT NANDA report coverage, Forbes, August 26, 2025
- MIT report: 95% of generative AI pilots at companies are failing, Fortune via Yahoo Finance, August 2025
- AI project failure rates are on the rise, CIO Dive on S&P Global Market Intelligence, March 2025
- 56% of CEOs see zero ROI from AI, Forbes, January 28, 2026
- Enterprise AI adoption in 2026, WRITER
- Anthropic Economic Index report, Anthropic, June 2026
- Klarna AI Customer Service: Replacing 700 Agents, Perspective AI
- AI Customer Service Cost Savings by Industry, Fin AI, 2026
