← Back to Insights

Insight

The Wrapper Was the Company

Ariel Agor
The Wrapper Was the Company

Listen · Read by Leo · click any word to jump

0:00 / · loading…

On July 9, 2026, OpenAI shipped ChatGPT Work. The pitch: give it a goal, get finished sheets, decks, docs, and small web apps back. GPT-5.6 in three tiers, Sol for flagship, Terra for balanced, Luna for cost. Bloomberg covered it the same day.

Around the same time, the Model Context Protocol team promoted its Enterprise-Managed Authorization extension to stable. Anthropic, Microsoft, Okta, and a growing list of MCP servers now sit behind a single identity provider gate. Asana, Atlassian, Canva, Figma, Linear, Supabase. One badge, every system.

Meanwhile Salesforce Agentforce, on the vendor side of the same trade, sits at $800 million ARR, up 169 percent year over year, 29,000 deals booked in a single quarter, 2.4 billion "agentic work units" processed in Slack, 85 percent of customer queries resolved without a human.

Somewhere between those three events, a chief technology officer at a mid-market company is being asked the same question they were asked in 2019: build or buy. The question is the same. The answer no longer exists. The framing is broken. Both sides of the trade collapsed inward, and what most CTOs are actually being sold in July 2026 is a claim on the shape of their work.

The Numbers Say Both Answers Are Wrong

In August 2025, MIT's Project NANDA released "The GenAI Divide: State of AI in Business 2025." One hundred and fifty leader interviews, three hundred and fifty employee surveys, three hundred public deployments. The headline: 95 percent of enterprise generative AI pilots delivered no measurable return.

Buried in the report, a subtler finding. Purchases from specialized vendors succeeded roughly 67 percent of the time. Internal builds succeeded about a third as often. The obvious read: buy. Everyone quotes that number now.

Gartner extended the picture in 2026. Eighty-nine percent of AI agent pilots never reach production. Forty percent of agentic AI projects will be canceled by the end of 2027. The cited causes are not model quality. They are cost overruns, unclear business value, and inadequate risk controls. Every one of those is a management gap, not an engineering gap.

Read the two together. Building loses because the pilot never ships. Buying wins twice as often but still fails one time in three. And the "wins" measured in the study included companies that later reversed course.

Klarna is the reversal case worth memorizing.

The Klarna Reversal

In February 2024, Klarna deployed an OpenAI-powered customer service bot globally. The company claimed the bot did the work of 700 human agents inside a month. Two point three million conversations. Resolution time down from 11 minutes to under 2. Forty million dollars in annualized profit.

By May 2025, CEO Sebastian Siemiatkowski told Bloomberg the company was hiring humans again. His quote: "As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality." Customer satisfaction had cratered. Complex cases, emotional cases, multi-step problems ate the bot alive.

By 2026 Klarna had rebuilt to a hybrid. AI on routine volume. Humans on edge cases and high-value accounts.

What did Klarna own during any of this? The CRM data, yes. The customer list. But the routing logic that decided which query hit the bot and which reached a human, the escalation criteria, the evaluation harness that should have caught the quality decline before the CSAT tanked, the trace log that would have shown where the failures clustered, none of that was theirs. The wrapper around the model belonged, in effect, to whoever supplied the model. Klarna had bought a shape they did not understand and could not audit.

Read the rehire as a receipt. The story sits underneath. Klarna did not know what the bot was doing until the customers told them, and by then the reputational damage had a nine-month lead on the fix.

What "Buy" Actually Ships in July 2026

The ChatGPT Work launch is the perfect case study of what "buy" means this month. The pitch has no shape. It takes a goal, breaks the goal into steps, walks across your connected apps for hours, and hands you back finished output.

Read the pitch as an economist would. The vendor sells the breakdown of your work into steps, the sequencing of those steps, the choice of which apps to touch in what order, the decision of when the job is done. Every one of those is a workflow decision that used to live inside your organization. Every one is now packaged inside the vendor.

The MCP Enterprise-Managed Authorization extension makes this trade cheaper for the buyer and more expensive to reverse. One SSO login, every internal server the identity provider approves. The convenience is real. So is the centralization. Once a single vendor holds the credentials to Asana, Atlassian, Canva, Figma, Linear, Salesforce, and Supabase for your organization, switching vendors is not a procurement decision. It is a systems reengineering project.

Salesforce's Agentforce numbers tell you what buying looks like from the vendor's side of the ledger. Twenty-nine thousand deals in one quarter. Nineteen trillion tokens processed lifetime. Twelve AI agents per customer on average, projected to grow 67 percent in two years. Every deal expands the surface area of the vendor's claim on the customer's workflow. Every token processed makes it harder to leave.

The customer, meanwhile, gets a dashboard. The dashboard shows resolution rates and cost savings. The dashboard does not show what the vendor's model will do differently in six months when GPT-6 or Claude Opus 5 lands and the routing behavior shifts under them.

The Real Question Is Not Build vs Buy

The debate that dominates every AI strategy consulting slide in July 2026 is the wrong debate. The framing pretends there are two options. There are five, at least, and they operate at different layers.

The model layer. You will rent this. Even if you host open weights on your own hardware, the reference weights come from Meta, DeepSeek, Mistral, or one of the frontier labs. DeepSeek V4 Flash matches OpenAI-class performance on several benchmarks now. Mistral Medium 3.5 introduced a reasoning intensity control. Llama 4 continues to grow. The model layer is a commodity being commoditized further every quarter. Do not own this. Even the labs cannot own their own weights against the next release.

The runtime layer. Where inference happens. AWS, Google Cloud, Azure, Oracle Cloud, Cloudflare, or your own machines. Portability matters more than price here. The Model Context Protocol makes the runtime layer more swappable in principle, but the EMA extension makes it stickier in practice. Design for swap.

The workflow layer. What the model does, in what order, for what business outcome. This is where value collects and where enterprises are most confused about ownership. ChatGPT Work sells this layer as a service. Salesforce Agentforce sells this layer as a platform. The MIT NANDA study found that the enterprises that succeeded were the ones that kept the workflow definition outside the vendor even when they bought the model and runtime.

The evaluation layer. How you know the model is behaving the way you think it is. This layer is invisible until it fails. Klarna did not have one. Ninety-five percent of the failed pilots in the NANDA study did not have one. This layer catches vendor drift, catches quality regression, catches the difference between the demo and the real workload. Own this or you are flying blind.

The interface layer. What the human sees. Buy this most of the time. Build it if the interface is your product.

The build-vs-buy question dissolves once you look at the layers. The correct answer for most enterprises in July 2026 goes like this. Rent the model. Host the runtime somewhere portable. Define the workflow yourself. Own the evaluation harness completely. Buy the interface unless the interface is the product.

That answer sits outside the build-vs-buy vocabulary. It is a discipline about which layer you refuse to let a vendor own.

When to Build vs Buy AI, Answered Properly

The Anglo-Saxon version of the question is: what do you own when the vendor changes?

Rent what depreciates. Own what compounds.

The model depreciates weekly. The runtime depreciates in months. The interface depreciates in years. Those are rentals. The workflow definition compounds every time you tune it for your customers. The evaluation harness compounds every time it catches a bad output. The trace log compounds every time you review a failure. Those are the assets.

Two examples to make this concrete.

A logistics company signs a contract with Salesforce Agentforce for customer service automation. The tempting move: point Agentforce at the CRM, tune the routing prompts in the Salesforce console, use Salesforce's dashboard for metrics. The wrapper sits inside the vendor. When Salesforce swaps model versions in Q4, the routing behavior shifts, the escalation rate drops, and the logistics company only notices when call volume to human agents jumps. They have no independent measure of the change because the measure lives inside the platform.

The disciplined move: Agentforce handles model calls, but the routing logic sits in a repository the customer owns. The escalation criteria live in a policy document versioned in git. The evaluation harness runs against a fixed test set of real historical tickets, nightly, and fires an alert when accuracy drops by more than two percent from the seven-day trailing average. The trace log ships to the customer's own observability stack. If Salesforce changes model behavior overnight, the customer sees the drift in a day. If Anthropic underprices Salesforce by 40 percent for the same workload next year, the customer can swap in a week.

A regional bank spins up an internal team to build a lending decision agent. The tempting move: fine-tune a base model on historical loan data, wrap a UI around it, deploy it to loan officers. Six months later the base model is obsolete, the fine-tuning is wasted, and the team's expertise is stuck in a snapshot of a moment.

The disciplined move: no fine-tuning. The reasoning happens on a rented frontier model with a strong prompt. The workflow definition lives in a directory the bank owns. The bank builds an evaluation harness that runs every candidate model against a golden set of historical decisions and compares outputs. When GPT-5.6 Sol wins on Tuesday and Claude Opus 5 wins on Wednesday, the bank swaps overnight. The bank's IP is the golden set, the evaluation, the workflow. The reasoning is rented.

Both of those enterprises look, from the outside, like they "bought" AI. Neither is building a foundation model. But both own the layers where value collects. Both survive vendor swaps. Both compound.

The failed pilots in the NANDA study, the canceled projects in the Gartner forecast, the Klarna reversal, all share a structure. The enterprise handed the workflow and the evaluation to a vendor and kept only the interface. When the vendor changed, the enterprise had no traction.

Distribution, Not Just Ownership

There is a second axis nobody puts on the strategy meeting whiteboard. Where does the model run, and what does it see?

The MCP EMA release makes cross-vendor authentication cheap. The DeepSeek and Mistral releases make cross-model swaps cheap. What does not get cheap is the distribution of your data across those vendors. Every time a workflow reaches into a new SaaS system through MCP, a copy of your operational data crosses a trust boundary.

The evaluation harness discipline covers model behavior. It does not cover data location. A serious wrapper in July 2026 includes a data-location policy. Which fields never leave the runtime you host. Which fields can hit the vendor's cache. Which fields must be redacted before crossing the MCP gate.

Klarna did not think about this. Neither did most of the 95 percent. The five percent that succeed built the policy before they bought the tool. Then they used the tool inside the policy.

What This Means For Your Next Six Months

If you are a CEO or founder trying to answer the build-vs-buy AI question this quarter, three concrete moves.

First, take an inventory of which layer of the AI stack each of your current pilots is owned at. Model, runtime, workflow, evaluation, interface. For each pilot, name the owner. If the workflow layer is owned by a vendor, treat that pilot as a rental, not an asset. Budget it that way.

Second, build an evaluation harness before you sign a vendor contract, not after. A frozen test set of your actual work. A nightly run. A threshold that fires an alert. This is the single highest-leverage piece of the wrapper. It costs a fraction of what a fine-tuning run costs. It is the difference between knowing your vendor drifted and hearing it from your customers.

Third, treat the model as a commodity even when it is not one yet. If your workflow depends on GPT-5.6 Sol's specific behavior on a specific edge case, you are locked in and you have not priced that lock-in. If your workflow depends on the presence of "a model that scores above X on your evaluation set," you are free.

These three moves are boring. They are not a moonshot. They are the difference between the 5 percent that make it and the 95 percent that hand a check to a vendor and wait for the ROI that never lands.

The Wrapper Is Not a Tool You Can Buy

Every serious enterprise AI story of the last two years has ended the same way. Klarna owned no wrapper. The NANDA failed pilots owned no wrapper. The canceled Gartner-tracked agentic projects owned no wrapper. The 5 percent that returned measurable value all owned a wrapper.

Nobody sells the wrapper. Not OpenAI, not Anthropic, not Salesforce, not any of the specialist vendors ChatGPT Work is being sold alongside. They cannot. If they sold the wrapper, they would sell you the thing that lets you leave them. The wrapper is the piece of the stack that makes the buy option safe and the build option worth doing. It has to come from inside your organization, or from a partner whose incentive is to build it for you rather than to lock you in.

This is why buying an off-the-shelf tool does not solve the AI question and why hiring a team to build a custom model rarely does either. What has to be designed is the wrapper. The evaluation, the workflow, the escalation policy, the trace log, the data-location rules, the vendor-swap discipline. That work is architecture. It is a strategic function. It has to be done deliberately by people who understand both the technology and the shape of your business, and it has to happen before the vendor contract is signed, not after.

That is where Agor AI Advisory comes in. We do not sell you a model. We do not sell you a chatbot. We work with you to design the wrapper that makes the model do useful work for your specific business, and continues to do useful work when the model underneath changes, because the model underneath will change, and the wrapper is what carries your investment forward.

Sources

Want this kind of automation working for your business?

Agor AI designs and ships the systems these posts describe, scoped in weeks, not quarters.

Book a Free Strategy Call