On September 11, Salesforce named seven agents. Casey handles service. Paige runs IT and HR. Carter shops. Hunter sells outbound. Marshall works supply chain. Piper runs inbound pipeline. Fin owns customer experience. The press release framed this as job-ready AI. The read for a buyer is different. Seven proper names, one for each of the roles where custom pilots have been dying quietly for two years. Read the move as a bet about buyers, not agents. Two failed custom pilots put most executives in a mood to take a fifty-fifty shot on a packaged agent rather than commission a third project that will not ship.
A year earlier, MIT's Project NANDA published a survey that put a number on how quiet that dying had been. Ninety-five percent of generative AI pilots produced no measurable return on the profit-and-loss statement. A hundred pilots start, five change a line on the income statement, and ninety-five are pinned to a corkboard somewhere in the innovation office. Gartner put a slower rhythm on the same idea in a June 25, 2025 press release: over forty percent of agentic AI projects would be cancelled by the end of 2027. Two research shops, one shape. The whole industry now calls this problem by the same phrase. Moving an AI pilot to production is what every consulting deck promises to help with, and every deck describes the same failure: pilot green, executives excited, then eighteen months of integration meetings where the green quietly turns brown.
The models are not the reason ninety-five out of a hundred pilots fail. GPT-4 could summarize a support ticket in 2023, and Claude Mythos 5.1 and OpenAI's newer families can now run tool calls over an enterprise's API surface with a level of reliability that would have looked like science fiction two years ago. The pilots pass. A pilot is scoped to pass. That is the whole problem.
What the Pilot Was Actually Measuring
A pilot is a controlled demonstration of capability. Someone picks a use case that is easy to describe to a steering committee. The support inbox. The invoice queue. The RFP response drafter. A small team assembles a test bench, feeds it a clean sample of documents, wires up a model, and lets it run against fifty carefully chosen examples. The output is scored by the same people who scoped the exercise. When the model gets forty-two out of fifty right, the pilot is called a success. A slide goes to the executive committee. A budget request follows. Nine months later the RFP drafter still lives on a laptop under a project manager's desk.
The reason this happens is not mysterious. The pilot was engineered to pass by narrowing scope until the model could not really fail. The fifty examples were the sunny path. The scoring was generous because the point of the exercise was to justify further investment, not to survive an audit. Nobody piloted the case where the CRM returned a null. Nobody piloted the vendor whose invoice arrived as a scanned photo of a fax. Nobody piloted the customer who asked the same question in Portuguese and English in the same message. Nobody piloted the day the model provider changed its default output format and every function call started returning unparseable JSON. The pilot could not fail because failure was outside the frame.
Production is the frame. Production is the CRM null, the faxed invoice, the bilingual message, the vendor's silent format change. Production is that thing plus a legal review, an information security review, a data residency review, an accessibility review, an integration test suite, a rollback plan, a runbook, an SLA, an on-call rotation, and a person whose name goes on the ticket when the model hallucinates a shipping address on a Sunday morning. The gap between the pilot and the production system is measured in ownership. Somebody's name has to go on the ticket at three in the morning, and the pilot never asked who.
Why Moving an AI Pilot to Production Kills Eighty-Eight Out of a Hundred
Iris.ai's September 2026 enterprise analysis put the pilot-to-production failure rate at eighty-eight percent. Astrafy tracked a broader denominator and reported that only about a third of pilots ever run in production. The numbers vary because "in production" is not a single line. Some pilots get called production the moment they start receiving live traffic on a single user's laptop. Some get called production only when the finance director sees the cost of the human team drop. The shape of the survey data is stable regardless. Somewhere between two thirds and nine tenths of the pilots that leave the demo room never become a business system.
The reasons are consistent across surveys. Nobody scoped the integration work. Nobody named the owner. Nobody wrote the ninety percent of the runbook that has nothing to do with the model. The steering committee that approved the pilot has moved on. The vendor who ran the demo is chasing a bigger logo. The internal champion changed roles. The regulator is still asking about the audit trail. The legal review is stuck on a clause about training data. The security team is refusing to whitelist the outbound domain. The finance team is asking how the monthly bill will scale if usage triples.
Any one of those could stop a project. In the average pilot, all of them arrive at once, at month seven, when the excitement has cooled and the executive sponsor has stopped answering the calendar invite. The pilot has done nothing wrong. The pilot has done exactly what a pilot was scoped to do. What the pilot did not do was prepare the organization to run a thing without a person carrying it every day. Nobody budgeted that work because that work does not photograph well in a board deck.
Salesforce read that failure pattern and made a bet. On September 11, 2026, Marc Benioff announced Agentforce 360 with the seven named agents and a governance layer called the AI Control Plane, designed to register, monitor, and govern agents across Salesforce and third-party platforms. Salesforce is wagering that a buyer who has been through two failed custom pilots will pay to be handed a Piper or a Hunter that already knows how to log its actions, obey a policy, and survive the security review. Buyers are increasingly receptive because the alternative is another eighteen-month integration project that ends with a slide.
The wager may pay off for Salesforce. The problem it does not solve for the buyer is that a packaged agent still lives inside the buyer's data model, the buyer's approval flows, the buyer's exception paths, and the buyer's compliance regime. Piper does not know that this company's inbound leads from Ohio hospitals get routed to a special enterprise team while everything else goes to SMB. That routing rule lives in the buyer's head or in a set of sales-ops macros nobody has read since 2022. The named agent solves the model problem. The organization still owns the integration problem, and the integration problem is what killed the last pilot.
Klarna, and What "Working" Looks Like
The Klarna story is the one every executive committee now has to answer for. In February 2024, Klarna's CEO Sebastian Siemiatkowski announced that its OpenAI-backed customer service agent had done the work of seven hundred human agents in its first month, cutting resolution time from eleven minutes to under two and projecting forty million dollars in annual savings. This was the strongest pilot-to-production story in enterprise AI. On May 8, 2025 the same CEO reversed course. Klarna started rehiring humans. Siemiatkowski's own framing was that Klarna had focused too much on efficiency and cost, and the result was lower quality, and that was not sustainable. By June 2026 the company described a hybrid model where AI handled routine work and humans handled complex and premium interactions, with cost per transaction down forty percent over two years.
Two things are true about that story. The pilot was real. The metrics that got called out in the original press release were not fabricated. AI did resolve tickets at speed. The production system had to be rewritten anyway. The version that worked in the demo did not work as the company's permanent customer service organization. The company kept the cost savings and gave back the arrogance of thinking the pilot had covered every case that mattered. Every executive weighing an AI workforce plan in 2026 now has to explain how their plan avoids the Klarna outcome. The honest answer is that Klarna did not do anything unusual. Klarna did what everyone does. The pilot passed. The production system failed. They rebuilt it.
The interesting number in the Klarna reversal is the forty percent cost per transaction reduction that survived the rebuild. That is the shape of a real production system. It is smaller than the pilot promised. It is durable. It shipped through a rewrite, and the rewrite was the point.
The Payroll Test
The framework I use with executive teams is a single question. Would you pay this agent what a person would earn to do this job. Not whether the demo works. Not whether you would buy the pilot at the ninety percent discount the vendor is offering. Would you write a payroll number next to this thing, tell the finance director the salary is loaded with fully-costed benefits, and let it show up on the org chart under someone's cost center. If the answer is yes, ship the thing. If the answer is no, keep piloting, and stop pretending the pilot is production work.
The payroll test is uncomfortable because it strips out the excitement about the model and forces a real number. The cost of running the agent at expected load, including inference, tool calls, human review time, incident response, retraining, and vendor lock-in premium. The value it produces in the same period, counted the way the finance director would count a human team's output. The reason ninety-five out of a hundred pilots fail the P&L test is that the pilot was never priced as a job. It was priced as an experiment. Experiments do not have to pay for themselves. Jobs do.
The 5 percent of pilots that succeeded, in MIT's telling, embedded memory and learning loops into high-value workflows and shipped tools that adapted with the organization. That is the technical version of the payroll test. The successful pilots crossed over because someone treated the second version as an operational hire, not a research toy. They built the memory. They wired the feedback. They budgeted the on-call. They named the owner. They wrote the runbook.
Rebuild the Second Version From Scratch
The single most useful thing an executive team can accept about moving an AI pilot to production is that the pilot code and the production code are not the same codebase. They are two different systems that share a use case. The pilot was optimized to pass a review. The production system is optimized to run a business function without a person carrying it every day. Ship the pilot as a memo, and kill it as soon as it works. Fund a second-version team that starts from scratch with different questions.
The different questions look like this. What breaks first when we double the load. What breaks first when we halve the quality of the input. What breaks first when the model provider changes its default output format. Who is on call the day it changes. What is the rollback plan when they are asleep. How do we know the agent is doing the wrong thing before the customer tells us. How do we retrain when the wrong thing is legal but bad. How does this thing get audited. What does the auditor need to see. How does that log get generated without hand-editing. Where does the log live. How long does it live. Who signs off on the retention policy.
None of those questions are model questions. All of them are organization questions. The organization is what is missing when a pilot passes and production stalls. The 5 percent that ship, the Klarna outcome that survived the rewrite, and the packaged Agentforce agents that Salesforce is now selling all sit on the same underlying insight. The moat is the operational discipline that surrounds the model, and the risk is that you shipped a pilot that could not carry that discipline because the pilot was never asked to.
What This Means for the Buyer This Quarter
Three practical shifts belong at the top of every executive AI review in the next quarter.
The first is a budget shift. Every pilot that gets a green light should be paired with a production commitment of at least three times the pilot budget, held in reserve, with the same executive sponsor's name on it. If the reserve does not exist, the pilot should not start. Piloting without funded production is the fastest way to end up in the ninety-five percent, and the pilot itself is often cheap enough that finance approves it without demanding the second budget. Approve them together, or approve neither.
The second is a scoping shift. Every pilot should be scoped against a real payroll number from day one. If the agent were a person doing this job, what would they be paid, what would the vendor charge, and how does the delta pay back the integration and operational cost. That number should live at the top of the project brief. Every progress review should read it aloud. When the number stops making sense, the project stops, or the scope changes to something where the number does make sense. This is the discipline the Klarna reversal produced. The cost per transaction reduction that stayed was the number the rebuild refused to negotiate away.
The third is a build shift. If the buyer intends to use packaged agents from Salesforce, ServiceNow, Microsoft, or the OpenAI enterprise family, the buyer still owns the integration architecture, the data model, the memory layer, and the governance surface those agents will run inside. Off-the-shelf named agents move the boundary of the integration problem. They do not remove it. A company that treats Agentforce as a plug-and-play solution has not read the eighty-eight percent number carefully enough. The named agent is the model. The integration is still yours.
The Real Work
The pilot was the easy part. The market has spent two years learning that. What comes next is the boring, expensive, uncoolly-photographed work of turning a model call into an operating capability. Budgeting for on-call rotations. Writing runbooks. Naming owners. Building memory. Wiring feedback loops. Instrumenting the boring twenty percent of edge cases that the pilot was allowed to skip. Publishing a payroll number and holding the project to it.
Leadership does that work, not tooling. The organizations that will show up in the next MIT survey with a P&L line to point at are the ones that decided, quarter by quarter, to fund the second version instead of the third pilot. The organizations that will show up in Gartner's cancellation number in 2027 are the ones that keep approving small pilots because they are cheap, and never approve the expensive production commitment that would actually change the income statement.
Architecting that transition is a design decision about how the organization intends to use AI. Either it is an experiment budget on the innovation ledger, or it is a set of operating capabilities on the P&L. Vendors will sell you named agents. Consultants will sell you frameworks. Neither of those is a substitute for a leadership team that has decided moving an AI pilot to production is a first-class business program with its own budget, owner, calendar, and standard of success. That decision cannot be outsourced.
Agor AI Advisory works with executive teams on exactly this transition. The second-version production build that survives the security review, the finance review, and the audit, and shows up as a number the CFO will underwrite. If the last eighteen months of AI experiments have produced a slide deck and a vague sense that something should be changing on the income statement, the payroll test is designed to break that pattern.
Sources
- Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, June 25, 2025
- MIT Finds 95% Of GenAI Pilots Fail Because Companies Avoid Friction, Forbes, August 26, 2025
- Klarna changes its AI tune and again recruits humans for customer service, CX Dive
- Klarna Reverses AI Push, Says Customers Prefer Human Support, Forbes, May 18, 2025
- Salesforce Makes Bold AI Play with Launch of Agentforce 360, TechRepublic
- Salesforce's Seven Named AI Agents, Kurums, September 2026
- MIT Report: 95% of Generative AI Pilots at Companies Are Failing, Yahoo Finance
- Enterprise AI Implementation: Why Pilots Never Scale, Perceptive Analytics
