On May 18, 2025, Klarna's chief executive Sebastian Siemiatkowski publicly said the company would start hiring humans again for customer service. This was the same executive who, fourteen months earlier, had compared his OpenAI powered chatbot to 700 full time support agents. The bot had handled 1.3 million conversations in its first month. It had cut average resolution time from eleven minutes to under two. He had said, on stage, that Klarna would eventually run itself with about half the people. Then, quietly, the customer satisfaction data got worse. The complex cases piled up in the queue where humans used to catch them. The savings the board had projected never fully arrived. So Klarna posted job listings for gig based support agents, and Siemiatkowski said, in as many words, that customers preferred talking to a person.
Read that as a piece of change management data. It is a signal about how AI adoption actually converges inside a business that ships to real customers every day.
The dominant playbook says the opposite. Executives commit, the workforce resists, the change managers overcome resistance with communication and training and quick wins, and eventually the organization catches up to the vision. Read any of the current AI adoption 2026 guides floating around and you will find the same arrow. Leadership on top. Skeptics on the bottom. The job of change management to bend the skeptics toward the vision.
At Klarna, the arrow ran the other way. The workforce Klarna had already laid off did not vote on the reversal. The customers did. The queue metrics did. The signal moved up from the operational floor, and the executive who had planted his flag on AI first had to walk it back in public.
At Duolingo, the same shape appeared in miniature. On April 28, 2025, chief executive Luis von Ahn issued an "AI first" memo. Employees would be evaluated on their AI use. Hiring would slow. Contractor headcount would fall. About a year later, on April 13, 2026, Fortune reported that Duolingo had dropped AI usage as a formal performance review metric. Von Ahn said employees had started asking "Do you just want us to use AI for AI's sake?" He conceded the question was reasonable.
Two of the most public AI first bets in the current cycle walked back the specific mechanic each one had bragged about. Both walked back for the same underlying reason. The people closest to the work had a more accurate model of where the tools broke than the executives who bought them.
The sightline is inverted
This is the piece the standard AI change management framework misses. It treats worker skepticism as a friction to overcome. It should be treating worker skepticism as the ground truth against which every executive prior gets updated.
Stanford's 2026 AI Index made the size of the miscalibration hard to argue with. Ninety seven percent of executives report benefiting from AI. Twenty nine percent see significant organizational return. Sixty one percent of executives say they trust AI for complex, business critical decisions. Nine percent of workers say the same. That is a fifty two point gap between the group buying the tools and the group operating them.
The standard reading of that gap is that workers are behind on the technology and need to catch up. The Klarna and Duolingo reversals suggest the more expensive reading is closer to true. Workers spend their day watching the model fail on the specific cases the demo did not cover. Executives spend their day inside decks that summarize the successful cases. In every prior enterprise software cycle, the executive priors were the informed ones. The board saw the roadmap; the intern saw the login screen. In this cycle, the interns and the customer service reps see the actual behavior of the system, and the board sees a curated slide.
That inversion is what makes AI change management a genuinely new discipline. The old digital-era playbook has the arrow pointing the wrong way. The old discipline was mostly about pulling people up to the level of a well understood system. The new discipline is mostly about pulling executive expectations down to the level of what the system actually does in production.
What MIT actually said
The MIT NANDA report that put the "ninety five percent of GenAI pilots deliver zero return" number into every board deck last summer buried the more useful finding in the middle. The five percent that worked were not the companies with the biggest AI budgets or the most senior AI leadership. They were the companies that had partnered with specialized vendors and, more importantly, had let line managers pick the workflows the AI would touch. Internal builds succeeded a third as often as vendor partnerships, and top down deployments succeeded rarely at all.
MIT also reported that more than half of GenAI spend went into sales and marketing tools, while the biggest measurable return came from back office automation. That is the same inversion in different clothes. The executives who controlled the budget put it where they read about AI. The operators knew where the actual pain sat. The budget went to the wrong workflows because nobody with a keyboard on the receiving end got a vote.
The Writer 2026 enterprise adoption study, which surveyed more than eight hundred enterprise leaders, put a version of this in numbers worth reading twice. Seventy nine percent of enterprises face ongoing AI challenges despite record investment. Only twenty nine percent see significant organizational return. The correlation between spend and outcome broke somewhere between 2024 and 2026, and it has not repaired.
Shadow AI is a signal, not a violation
On July 7, 2026, Smarsh and FTI Consulting published the results of their 2026 Enterprise AI Trends Study. Fifty five percent of enterprises are actively deploying AI. Twenty six percent say their governance frameworks keep pace with deployment. Thirty percent have the ability to detect shadow AI, meaning the tools employees are already using without approval.
Read the shadow AI number the way most compliance briefings read it and you get a story about risk exposure and rogue employees. Read it as change management data and you get a different story. The seventy percent of enterprises that cannot see shadow AI cannot see the actual working diffusion of AI inside their walls. They have committed a large budget to a governed program that is smaller than the ungoverned use their people are already running.
The workforce built a shadow adoption pattern because the sanctioned pattern was worse. Nobody uses a personal ChatGPT account at work because they hate the enterprise Copilot license the IT team just bought them. They do it because the enterprise Copilot license, in the workflow they actually have, does not do the thing they need. The shadow use is a running experiment about which AI wraps well around real work. Governance that treats it as a violation to be shut down is throwing away the most useful adoption dataset the company will ever collect.
The playbook needs a new arrow
If you accept the inversion, AI change management stops looking like a communications problem and starts looking like a sensor problem. The organization has to build a way for the ground truth to travel upward faster than the enthusiasm travels downward. Without that, you get more Klarnas.
A few concrete moves fall out of this framing. None of them look like the standard AI adoption slide.
The first move is to invert who owns the pilot list. In most enterprises, the AI center of excellence picks candidate workflows, ranks them, and imposes the winners on the operating teams. Flip that. Let operators nominate the workflows they think an AI can wrap around, and let them rank the current sanctioned tools against the shadow tools they are already trying. The center of excellence becomes an accelerator for the workflows the front line has already validated on its own time, rather than an imposer of workflows the front line has never asked for.
The second move is to publish the failure log with the wins. Every AI deployment ships two kinds of results. Cases the model handled well and cases the model handled badly. The dominant reporting pattern shows the board a resolution rate and buries the failure taxonomy in an appendix nobody reads. Publish both, at the same fidelity, on the same page, on a cadence the board actually sees. The board should read the top ten open failure modes before it reads the aggregate resolution rate. Klarna's reversal happened because a failure mode (complex cases dying in a bot queue that had no human backstop) accumulated to the point that the resolution rate could no longer hide it. If the failure taxonomy had been on the board's monthly page, the reversal would have arrived a year earlier and cost eight figures less.
The third move is to hire for the seam, not the tool. Most AI role postings in 2026 look for a prompt engineer or a model operations specialist or a machine learning platform lead. The scarcer role is the person who can sit with a customer support agent, a claims adjuster, a warehouse manager, and rewrite the workflow so that the AI takes the cases it can close and hands off the rest with the context intact. That role has almost no formal training pipeline. It is where the actual productivity gains from AI are getting locked up right now, and it is the role that almost no AI change management program is explicitly building. The Klarna hybrid model that finally worked, meaning AI for routine cases and humans for escalations with the handoff engineered rather than assumed, required exactly this kind of seam design. Klarna paid for it by rehiring, having already fired for it, then discovering that the seam was where the value lived.
The fourth move is to treat every rollback as a first class deliverable. In the current cycle, when a team pulls an AI feature back to a human because the AI could not carry the case, the rollback gets written up as an embarrassment or, more often, not written up at all. Duolingo pulled the AI mandate. Klarna pulled the customer service layoff plan. Both companies now own more accurate information about where AI fits inside their business than any of their peers who never made a public reversal. That information is the asset. The reversal is the paid tuition. If you cannot afford to publish your rollbacks internally with the same status as your launches, you are asking every future project to relearn the tuition at full price.
Why this cannot come from a vendor
Every one of these moves is architectural rather than instrumental. You cannot buy a piece of software that inverts your pilot ownership, publishes your failure log to the board, hires the seam designer, or gives rollback the same status as launch. Those are decisions about how the enterprise itself is wired. A vendor can sell you a copilot, an agent runtime, a governance dashboard, a training curriculum. A vendor cannot sell you an organization where the failure modes reach the board on the same page as the wins. That has to be built into how the company runs.
This is where the AI change management category, as most consultancies sell it, is misspecified. The default offering wraps AI adoption in the same posture the big firms sold for cloud migrations and ERP rollouts. Executive alignment sessions. Communications cascades. Persona based training. Champion networks. That posture assumes the executive vision is the accurate model of the destination and the workforce is the noise to be filtered. In AI, the executive vision is the noisy model. The workforce is closer to the signal. Every deliverable in the standard change management stack is optimized for the wrong direction of information flow.
The companies that are quietly doing well with AI in 2026 have almost all done the same architectural thing. They have inverted the reporting so that operator observations travel to the board at the same speed as vendor demos, and they have given operators the authority to pick which workflows the AI touches. They are not sending fewer memos. They are receiving more of them, from the people who actually watch the tools work and fail.
The board's job is to update, not to hold the line
There is a version of executive courage that says: hold the line, trust the roadmap, do not let the noise from the front lines derail the transformation. That version was correct for many prior technology cycles. It is wrong for this one.
The AI stack is moving faster than any prior stack. The failure modes surface faster and change more often. Every model release rewrites which cases now work and which now break. Holding the line on a roadmap set in Q1 of 2026, on a model that has been superseded twice since, using an integration pattern the workforce has already routed around with a shadow tool, is not courage. It is a slow reversal that will eventually happen anyway, on worse terms, after the tuition has been paid twice.
The braver move for a board in 2026 is to publicly update. To say, in a shareholder letter, that the pilot the company launched last year has been rebuilt because the operators found a better shape. To reframe the reversal as evidence that the internal sensor is working. To reward the manager who caught the failure mode early instead of the manager who defended the original plan. That posture is uncomfortable because the incentives inside most large companies penalize public updating. AI change management, done seriously, is the work of rebuilding those incentives so that updating is the norm and holding the line is the anomaly.
What to do about this on Monday
The change management theme people are searching for right now is not really about how to communicate an AI rollout. It is about how to run a company where the AI ground truth lives closer to the floor than to the ceiling, and the reporting has to be rebuilt around that fact.
This is a design job. It touches how workflows are chosen, how failures are surfaced, how rollbacks are treated, how compensation rewards early updating, how governance handles shadow use, and how the board reads the monthly page. It is the exact work that off the shelf tools cannot ship and that the big consultancy playbooks are structurally wrong about, because their offering was built for the last inversion, not this one.
Agor AI Advisory is built for this exact problem. We come in as architects of the reporting and decision loops that let AI ground truth reach the people making the calls, so the reversal that would have arrived on the board's second anniversary of the program arrives on its second month instead. We rewire the pilot selection, the failure taxonomy, the seam design, and the incentive structure so that your AI program stops paying tuition on the same lesson twice. Every executive team that has read the Klarna and Duolingo stories and quietly wondered whether their own program has the same failure mode waiting has an answer here. Fix the direction of the arrow before the reversal writes itself into your press cycle.
Sources
- Klarna Reverses AI Push, Says Customers Prefer Human Support, Forbes, May 18, 2025
- Klarna Is Hiring Customer Service Agents After AI Couldn't Cut It, Entrepreneur
- Duolingo CEO backs off from evaluating employees on AI usage, Fortune, April 13, 2026
- MIT report: 95% of generative AI pilots at companies are failing, AOL Finance
- Enterprise AI adoption in 2026: Why 79% face challenges despite high investment, Writer
- Stanford 2026 AI Index: What Business Leaders Need to Know, SAPInsider
- New Smarsh Research Finds Enterprises Are Deploying AI Faster Than They Can Govern It, Smarsh, July 7, 2026
