← Back to Insights

Insight

The Ledger Was Wrong

Ariel Agor
The Ledger Was Wrong

Listen · Read by Leo · click any word to jump

0:00 / · loading…

On July 22, 2025, MIT's Project NANDA published a report titled "The GenAI Divide: State of AI in Business 2025." It got quoted in enterprise IT press for a week and then went quiet. Through August and September 2026 it came back everywhere. MarketScale, Healthcare IT News, CIO Dive, Legal.io, half a dozen advisory firms. The finding is one sentence long. Ninety-five percent of generative AI pilots at major enterprises have delivered no measurable P&L impact. Enterprise spend on generative AI nearly tripled in 2025 to around $37 billion. Ninety-five percent of that spend produced nothing a CFO could point at on the books.

The number is not the problem. The number is real, and it has been reproduced by every serious survey since. S&P Global Market Intelligence found in October 2025 that 42 percent of companies scrapped most of their AI initiatives that year, up from 17 percent the year before. IBM put the share of initiatives delivering expected ROI at 25 percent. Morgan Stanley found that 21 percent of S&P 500 companies could cite a measurable AI benefit at all. Different surveys, different populations, roughly the same distribution.

The problem is what boards are doing with the number. They are reading it as a signal to slow down, cut the pilot budget, wait for the technology to mature. Every one of those moves is a category error. The report says the pilots produced no P&L impact. It does not say the pilots did nothing. It says the books had no place to record what happened.

Measuring ROI on AI initiatives is a discipline nobody built

On March 24, 2026, Gartner published a research note aimed at CFOs. Only 29 percent of executives can confidently measure AI ROI today. Seventy-nine percent already see productivity gains. The two numbers do not reconcile until you look at what they are measuring. Productivity is real inside the workflow. The measurement infrastructure to prove it on the P&L is not built. Forty-four percent of organizations have adopted any form of AI FinOps practice at all. The rest are running an untracked spend against an untracked return and calling the gap between them "no ROI."

McKinsey caught the same pattern from the other side. Only 21 percent of generative AI adopters fundamentally rebuilt the workflow the AI was supposed to run. That group is 3.6 times more likely to see EBIT impact above 5 percent. Ninety-five percent of pilots produce nothing because 79 percent of buyers did nothing to the process the pilot was supposed to change. They wired the model into the current flow, left every human step in place, and waited for a number to move. Nothing moved. Nothing had been asked to.

Measuring ROI on AI initiatives is not a spreadsheet exercise you run after a pilot ships. It is an architectural exercise you run before the pilot ships, because the only ROI a model can produce is the difference between a workflow before it was rebuilt and the same workflow after. If the workflow was not rebuilt, there is no difference to measure. If it was rebuilt and no baseline was captured, the difference is impossible to reconstruct. In both cases the CFO gets the same answer at the six-month review, which is silence, which they read as failure.

The reference cases everyone quotes and nobody copies

Klarna is the case that turns up in every slide deck. In the first month after launch, its AI customer service assistant handled 2.3 million conversations. That is the work of roughly 700 full-time agents. Customer satisfaction held. Average resolution time fell from 11 minutes to under 2. On a headcount basis the return is unambiguous. On the P&L, it is buried across three lines. Support cost of goods sold gets smaller. The refund cycle gets shorter, which lowers working capital tied up in disputed transactions. New agent hiring slows against rising volume. Every one of those lines has other drivers. A CFO looking for a single row labeled "AI ROI" will not find it, because Klarna did not book one. What Klarna did was rewire the customer service function around the model, then keep counting the numbers the CFO was already counting.

Duolingo is the second reference case, and the more instructive one. In Q1 2026 the company published 20,500 new course units. The quarterly average in 2025 was 7,100. In 2024 it was 1,800. The mechanism was AI-assisted content generation, which lets a smaller team ship an order of magnitude more product. The return did not land in an AI budget line. It landed in a rising top line, a growing engagement number, and 56.5 million daily active users in Q1, up 21 percent year over year. CEO Luis von Ahn publicly stopped measuring employees' AI usage in performance reviews earlier this year, after staff pushed back on an "AI-first" internal memo. What he kept measuring was course throughput, retention, and paid conversions. The tool disappears from the ledger. The output shows up in the earnings call.

The pattern in both cases is the same. The company did not add a new row to the P&L. It changed what the existing rows meant. That is the shape of a real return.

The pilot that hides its own return

Every measurement failure I have seen in the last twelve months traces to the same shape. A CIO gets budget for an AI pilot. The pilot runs alongside an existing process. The process does not change. When finance asks for ROI at the six-month review, there is no baseline to compare against, because the baseline was the process the pilot was supposed to replace, and the process is still there, running in parallel. The pilot is adding cost, producing nothing the books can distinguish from noise, and the natural read is that the pilot failed. What actually failed is the decision to run the pilot alongside the process instead of through it.

The 5 percent that gets a return does the opposite. They pick one workflow. They take the whole loop apart. They put it back together with the model doing the throughput-carrying steps. They measure the new loop against the old one. That comparison is cheap to set up if you do it before you deploy and expensive to reconstruct after. Most companies discover they should have captured the baseline six months after they lost the chance. This is the single most common regret I hear from operating executives on a second consultation. Nobody tracked the before.

What actually goes on the line

There are five categories of return that show up when you build the measurement infrastructure alongside the workflow rebuild rather than after it. None of them require a new P&L row. All of them require a baseline captured on day zero.

Cycle time on a real workflow

Klarna's 11 minutes to 2 minutes is not a productivity metric. It is a working-capital metric. Every minute of shorter resolution time compresses the window in which a disputed transaction sits on the balance sheet as a contingent liability. Every minute compresses the customer's decision window between "I paid Klarna" and "I bought again." That compression shows up in trailing-twelve-month conversion and in average receivables days. The rebuild is measured in seconds. The return is measured in dollars per quarter. The trick is knowing which seconds to measure before the rebuild starts.

Deflected headcount growth

The number finance recognizes is layoffs. The number that matters is the hire that never happens. A support organization growing at 20 percent per year, backed by AI-augmented agents each carrying three times the ticket load, does not fire anyone. It slows the hiring plan. The right way to book this return is to publish the pre-rebuild hiring plan alongside the rebuild, then measure actual hires against it every quarter. The delta is real money. It never shows up in any accounting system unless somebody puts it there deliberately.

Quality captured as revenue

Duolingo's course output went up more than 3 times. The revenue effect is subtle and slow. More course inventory raises the odds a new user finds a language pair that fits, which lifts the retention curve, which lifts lifetime value, which lifts the price the company can charge for the paid tier. Every link in that chain is measurable. Nobody measures the whole chain unless somebody commits to reporting it end to end. Most companies stop at "we shipped more content" and lose the compounding.

Optionality that gets exercised

An enterprise that rebuilds one workflow with AI learns how to rebuild the next one in a fraction of the time. The second rebuild is worth more than the first, because the muscle is now there. The third is worth more again. This is a real capital effect and it is invisible to first-year ROI calculations. The way to book it is a running tally of workflows rebuilt per quarter, with each rebuild's incremental cost tracked against the previous one's. If the cost per rebuild is falling, the balance sheet has a real asset called institutional AI capacity. If it is flat or rising, the pilots are staying artisanal and none of them is compounding.

The rebuild itself as capital

The most expensive category and the most under-tracked. Every workflow rebuild produces a durable artifact. A rewired customer service tree. A revised content pipeline. A restructured underwriting flow. These are capital in the accounting sense. They are amortizable across future revenue. A company that rebuilds ten workflows in 2026 has ten pieces of process capital on the books at the end of the year that were not there at the start. Nobody books them. Nobody amortizes them. They get expensed as consulting fees and forgotten. The 5 percent that measures ROI properly treats each rebuild as a capitalized project with a service life. The 95 percent that gets zero measurable return has treated the whole spend as opex against a P&L that is measuring the wrong period.

Why the vendor cannot fix this

Every AI vendor selling into the enterprise right now is optimizing for tokens consumed, seats deployed, or requests per minute. None of those metrics maps to any of the five categories above. The vendor cannot build the ROI ledger for you, because the vendor does not know what the pre-rebuild baseline was, or which workflow you rewrote, or which hire you did not make. The ledger has to be constructed inside the company by somebody with access to the pre-rebuild books and the post-rebuild operations. That person is either the CFO's office, the COO's office, or an outside advisory that sits between them.

The default failure mode is to hand the measurement job to the vendor's customer success team. They report tokens, session counts, and satisfaction scores. Six months later the CFO asks where the P&L impact is and the customer success team hands over a deck of usage graphs. The graphs are honest. They also do not answer the question. The vendor was never in a position to answer it.

Gartner made this same point in a different register earlier this year. Forty-five percent of CFOs are spending their AI budget on the wrong outcomes. AI platform spend hit $64 billion in 2026 and rose more than 35 percent year over year while only 44 percent of buying organizations had any FinOps practice around it. The money is real. The measurement is missing. The vendor is happy to invoice against the missing measurement forever.

The architecture the number is asking for

The 95 percent statistic, read correctly, is a design brief. It says the current shape of enterprise AI deployment produces no visible return because it lacks three things the 5 percent has. A workflow rebuild before deployment. A baseline captured before the rebuild. A measurement infrastructure that reports outcomes in the categories the P&L already recognizes.

Building those three things is not a tool purchase. It is architecture work. It requires deciding which workflow to rebuild first, which parts of the loop the model owns, which parts the humans still own, how the outputs get measured, which existing accounting rows will change, and how the rebuild gets capitalized. None of that is inside any vendor's product. All of it is the work Agor AI Advisory does with clients before a single line of production code ships.

Every board and CFO in the country right now is being handed a version of the 95 percent number and being asked to draw a conclusion. Drawing the wrong one costs a year, because the reflex to slow down means the workflow rebuilds never get architected, which means the pilots keep failing for the reason they were already failing, which produces another year of zero measurable return, which validates the reflex, which slows spending further. That loop compounds the wrong way for as long as the ledger stays wrong.

The companies that will own their industries in 2028 are the ones that read the 95 percent number in September 2026 as an architectural indictment, and started rebuilding the workflow before they scaled the pilot. If your organization is one of them, the work begins with a serious conversation about which workflow, which baseline, which ledger, and which capital treatment. That is the conversation Agor AI Advisory exists to run.

Sources

Want this kind of automation working for your business?

Agor AI designs and ships the systems these posts describe, scoped in weeks, not quarters.

Book a Free Strategy Call