← Back to Insights

Insight

The Model Was the Cheap Part

Ariel Agor
The Model Was the Cheap Part

Listen · Read by Leo · click any word to jump

0:00 / · loading…

On July 2, 2026, Microsoft announced the Frontier Company. The unit ships with $2.5 billion in funding and 6,000 engineers. Judson Althoff, CEO of Microsoft Commercial Business, was blunt about the role. These engineers will sit inside customer buildings, wear customer badges, and stand up production AI systems next to customer staff.

The move fit a pattern that started building six months earlier. Amazon Web Services committed roughly $1 billion to a similar unit. On May 11, 2026, OpenAI formalized its own version and named it The Deployment Company, a majority-owned joint venture built to embed OpenAI engineers inside enterprise customers. Anthropic scaled its Applied AI Engineer program past 100 people. Cursor stood up a field team of its own. Accenture launched a dedicated Microsoft Forward Deployed practice on the same July 2 announcement. Capgemini, KPMG, EY, and PwC signed up as delivery partners.

Every major AI vendor now runs a staffing arm.

The framing is polite. The confession is loud. The models are done. The bottleneck is your company.

The number no one is proud of

Read one number from the last twelve months and the rest of the market makes sense. MIT's Project NANDA published a report called The GenAI Divide. Ninety-five percent of enterprise generative AI pilots produced no measurable effect on profit and loss. Roughly five percent extracted real value. The remainder sat in Confluence pages as evidence of intent.

The report went viral in the summer of 2025. Trade press treated it as a scandal. The vendors treated it as a specification. Every forward-deployed engineering unit stood up since is a direct answer to that number. If the models cannot ship themselves into revenue, we will send bodies to make them ship.

That is a strategic disclosure worth reading twice. The people who build the models believe the models are ready. What they no longer believe is that the market can install them alone. So they are stapling implementation labor to the software and calling it a product.

What a generative AI business use case actually costs

The phrase "generative AI business use case" reads like a menu item. Executives use it in board decks. Vendors use it in pitch pages. The word "use case" implies a bounded thing with a price tag. Pick one. Deploy it. Book the return.

The MIT number tells you where the price tag actually lives. The model is a small line. The engineering to wire the model into a workflow that pays the bill is the whole rest of the sheet. Every generative AI business use case that produced a real number in the last year has the same shape. Someone rebuilt the process around the model. Someone rewrote the data contract. Someone changed who signs what. Someone rewrote the incentives inside the team so the humans use the output instead of ignoring it.

That work is a rebuild, and rebuilds do not fit inside a purchase order.

Which is exactly what Microsoft, AWS, Anthropic, and OpenAI are quietly telling you when they mail you their engineers. The tool alone will not move your P&L. The rebuild does that job. If your team will not run the rebuild, they will run it for you, on their clock, at their rate.

The three roads a CEO can take right now

You have three doors. Every enterprise leader reading this has already picked one, whether they know it or not.

Door one is the pilot. Buy the tool, launch the working group, put six months on the calendar, produce a case study. This is the ninety-five percent path. The failure lives in what did not get touched. Nothing else changes, so nothing else pays.

Door two is the vendor's engineer. Take the FDE offer. Microsoft sends people. AWS sends people. Anthropic sends people. They wire the model into your workflows, publish a shared runbook, and hand you a system that runs. You get a working use case. You also get a working use case that keeps working only while the vendor's engineer keeps working there. The seams belong to them. The retraining discipline belongs to them. The kill switch belongs to them. The model can be swapped tomorrow and you will pay the swap cost again. Your line item shrank. Your dependency graph tripled.

Door three is the one almost nobody takes. You architect the integration yourself. You own the seams between the model and your data. You own the evaluation harness that says the output is good enough. You own the rollback path. You own the retraining loop. You buy the model as a component. You buy FDE hours as tactical labor, capped and scoped. You never let the vendor own the shape of your process.

Door three is harder to sell to a CFO. It is also the only door that yields compounding advantage. Every hour spent inside door three raises the value of every subsequent generative AI business use case you deploy, because the seams already exist and the next model plugs in behind them.

The seam is where the value hides

Look at the five percent of companies MIT called "learners." Read the case studies from Anthropic's Applied AI group, the OpenAI Solutions team, and Cursor's field engineers. The pattern is consistent and it does not point at the model.

Klarna cited its AI customer service agent in 2024 as doing the work of 700 human agents. By 2025, the company was hiring humans back to handle harder queries. The number that stayed positive was the reworked escalation ladder, the CRM data contract, and the QA scorecard the human agents used to correct the model. The model was one moving part inside a rebuilt operation. The rebuild carried the P&L, and the rebuild is what Klarna owned.

Bloomberg's terminal has been shipping retrieval-augmented workflows to fixed-income desks for over a year. The value gets locked in by the research asset, the query schema, and how analyst usage feeds the next iteration. The model rotates. The library stays. The library is the enterprise asset.

Even Cursor, the fastest-growing developer tool on the market and now the operator of its own FDE arm, is public about what actually gets billed. The model comes free with the seat. The wiring into a customer's specific repo layout, review conventions, and release process is what carries the invoice.

The seam is the strategy. Once you accept that, the vendor-versus-build argument dissolves. You are neither building nor buying a model. You are building or buying seams.

Why the current market wants you inside door two

The FDE model is elegant for the vendor. It solves three commercial problems at once.

It converts commoditized inference into a services contract with a much higher take rate. Microsoft charges you for tokens on the Azure meter and charges you for hours of Frontier engineer on a services line at the same time. Tokens are elastic and cheap. Hours are sticky and expensive.

It locks the seams into the vendor's platform. Every runbook the Frontier engineer writes cites Azure services, uses Azure identity, and stores state in Azure blobs. The runbook is portable in the way a house is portable if you own the foundation but not the land.

It gives the vendor a live feedback loop from real customer data at industrial scale. When Anthropic's Applied AI team writes a workflow for a bank, the anonymized learnings feed the next Claude release. When Microsoft's Frontier engineer writes a workflow for a manufacturer, Copilot for Manufacturing ships six months later with the exact pattern baked in. The customer paid to train the vendor's next product.

None of this is malicious. It is normal commerce. The vendor is optimizing a business. You are the market. Understanding the geometry is your job, not theirs.

The procurement discipline nobody has yet

Corporate procurement teams are still writing generative AI contracts as if they were buying SaaS seats or consulting hours. Neither template survives contact with a modern FDE deal.

A SaaS seat is a stable object. Every user gets the same features, priced per head, capped by concurrency. Generative AI does not obey that shape. The same seat this quarter costs you three times as much next quarter if the model gets more expensive, if the agent takes more turns, if the reasoning chain grows longer. Token spend is a variable, and your CFO has no line for a variable.

A consulting hour is also a stable object. You buy a body for a period, you supervise the body, and you own the deliverable. FDE hours are billed like consulting but produce artifacts that only work inside the vendor's platform. You paid for a body, and the body left a house built on rented land.

The right procurement discipline for door three is a three-line contract shape. First, a metered model tier with an exit clause tied to portability tests you own. If the vendor cannot show a working substitution against your golden evaluation set inside ten business days, you have grounds to leave. Second, an FDE hour cap tied to specific architectural artifacts you keep on your side. Every hour produces a spec, a data contract, or a policy definition that lives in your repo, not theirs. Third, a data reciprocity clause that governs what learnings from your workflows can and cannot feed the vendor's next product.

Almost no one writes contracts this way yet. The vendors will not offer this shape. You have to bring it.

What ownership looks like at the seam layer

Owning the seams does not mean hiring a large in-house AI team. That was the 2019 playbook. It failed because the models moved faster than the teams could specialize.

The current shape is different. A single architect and a small integration group can own the seams for a mid-sized company if the architecture is right. What the architect owns is not a model. It is a spine.

The data plane

Every prompt that leaves your building, every response that comes back, and every tool call the model made in between gets logged, replayed, and evaluated against a golden set that lives in your repo. When the vendor releases a new model, and in July 2026 alone that meant Claude Opus 5 on July 24, three new Qwen models, Kimi K3, Gemini 3.6 Flash, poolside's Laguna S 2.1, Ling-3.0-flash, and FLUX 3 across seven days, you rerun the golden set overnight and know within twelve hours whether the new model is better, worse, or sideways for your workflows. No PowerPoint. Just numbers.

The identity plane

Every agent that acts on your behalf gets a scoped token, a rate limit, and an audit log. When a model asks to send an email, publish a change, or draw down a budget, a policy layer decides whether the action is allowed and a receipt is written. The receipt outlives the model, the vendor, and the contract. If a regulator, an auditor, or a customer asks who did what and why, the answer sits in your storage account, not the vendor's.

The workflow plane

Humans and agents share a queue. Work moves through named stages. Each stage has an owner, a definition of done, and an escape hatch back to a human. When a model changes, the stages do not. When a process changes, the stages evolve on a schedule you set, not on the schedule of whichever vendor released a new model last Tuesday.

Own those three planes and you can rent any model. Rent any FDE team. Buy any tool. Nothing about a vendor decision is existential because none of the vendor decisions live inside your spine.

The August 3 window

The reason to think about this in early August, and not next quarter, is that the window in which door three is defensible is short.

Every FDE contract signed in the next six months writes another set of seams into the vendor's platform. Every runbook written by a Frontier engineer becomes another piece of institutional knowledge you cannot lift out. Every dashboard built inside Microsoft Fabric, Amazon Bedrock, or Google Vertex is another dependency you will pay to migrate someday.

The MIT number will fall. The five percent will grow. Not because the average enterprise got smarter. Because the vendors sent enough engineers to grind results out one customer at a time. The results will be real. The rent will be permanent.

If you architect your own spine now, the value the FDE units create shows up on your books. If you outsource the architecture, the value shows up on the vendor's books and you rent it back forever.

That is what "generative AI business use cases" means in the second half of 2026. The question stopped being which use case to fund. The question is who owns the seams when the use case works.

Architect this, do not buy it

The vendors will tell you their engineers are your engineers. They are their engineers on loan, and they are excellent. They will help you ship, and they will leave a system behind that only they can maintain. That is a defensible service. It is a terrible foundation for your strategy.

Buying an off-the-shelf AI capability at this moment is like renting your kitchen from the restaurant next door. Dinner shows up. The kitchen goes home at midnight. You never learn to cook. Next quarter, the price goes up.

Architecting your own spine is harder to sell in a board deck, so the pattern that wins is the one where a small outside team helps you design the spine, teaches your architect to own it, and leaves. The outside team is a catalyst. When they leave, you keep the leverage. The compounding sits on your books.

That is the work Agor AI Advisory does. We design and stand up the data, identity, and workflow planes that let you rent models freely without renting your future. We move fast because we have shipped this shape before. We are gone when your team owns it. What stays behind is a spine that pays down every month, no matter which vendor's badge is worn in the lobby next quarter.

Sources

The three doors: pilot, vendor FDE, own the seams

The post claims a CEO has exactly three options and that only door three compounds, but the tradeoffs are scattered across four sections of prose. Fifteen seconds with this table shows the reader that what separates the doors is not cost or speed but who owns the seams when the use case works.

  • 95% of enterprise generative AI pilots moved no P&L. That number is a specification, not a scandal — every forward-deployed engineering unit stood up since is an answer to it.
  • You are not choosing between building and buying a model. You are choosing between building and buying seams.
  • Door two shrinks your line item and triples your dependency graph. The results are real. The rent is permanent.
Who does the rebuildWho owns the seamsWhat compounds
Door 1 — The pilotCheapest to approve, and the only door with a documented failure rate attached to it.Nobody. The tool arrives, the process does not change.Nothing to own. No seams were built.Nothing. This is the 95% path MIT measured.
Door 2 — The vendor's engineerYou get a working use case fast; it keeps working only while their engineer keeps working there.Microsoft, AWS, Anthropic, or OpenAI engineers, inside your building, on their clock.The vendor. Runbooks cite their services, their identity, their storage.Their platform. Your workflow patterns ship back as their next product feature.
Door 3 — Own the spineHardest to sell to a CFO, because the line item is architecture rather than a deliverable.Your architect plus a small integration group; FDE hours bought capped and scoped.You. Data plane, identity plane, and workflow plane live in your repo.Every subsequent use case. The seams already exist and the next model plugs in behind them.

Source: The three doors section of this post, with the dependency and ownership claims drawn from the post's own analysis of Microsoft Frontier ($2.5B / 6,000 engineers, July 2, 2026), the MIT Project NANDA 95% pilot-failure figure, and the seam-layer argument in the same piece. · verified · as of 2026-08-03

Want this kind of automation working for your business?

Agor AI designs and ships the systems these posts describe, scoped in weeks, not quarters.

Book a Free Strategy Call