← Back to Insights

Insight

Thinking Sits in COGS

Ariel Agor
Thinking Sits in COGS

Listen · Read by Leo · click any word to jump

0:00 / · loading…

On August 17, 2026, Gartner published a research note with a phrase that will outlive the news cycle. They call it the Inference Paradox. The model gets cheaper. The bill gets bigger. Per-token prices are down roughly 95 percent since 2022 and heading lower still. Per-workflow inference costs for agentic systems are on track to rise more than fivefold through 2028. Both curves are real. Both are happening at once.

That is the whole story of the economics of AI agents in one sentence. Cheaper tokens buy more expensive tasks. The reasoning loop swallows every gain the model curve delivers. Every C-level operator making an AI decision this quarter is buying inside that paradox, whether they can name it or not.

For a decade, software companies grew fat on a simple structural fact. Once the product shipped, the marginal cost of serving another customer was near zero. Storage was cheap. Bandwidth was cheap. Compute for a page render was cheap. Cost of goods sold sat somewhere around ten to twenty percent of revenue and gross margins ran above eighty. That number, that gross margin, is what SaaS valuations were built on. Rule of Forty math assumes it. Net-dollar retention assumes it. Every LTV-to-CAC ratio the last two decades of enterprise investors quoted assumes it.

That gross margin no longer exists at AI-native companies. Bessemer's February 2026 pricing playbook and a16z's follow-up work both put AI-first gross margins in the 50 to 60 percent range. ICONIQ's 2026 growth report shows inference alone consuming about 23 percent of revenue at scaling-stage AI B2B firms. The old SaaS ceiling dropped by thirty points. The reason is a single new line item that never used to exist. Every reasoning step your customer executes costs you real cash to somebody else's GPU cluster.

The Line Item That Wasn't There

When Salesforce shipped a CRM record write in 2018, the marginal cost to Salesforce was somewhere south of a hundredth of a cent. The customer paid a seat license. The seat license paid for people, R&D, sales commissions, and a thin slice for AWS. Nothing about the customer's next click meaningfully changed Salesforce's bill.

When Salesforce Agentforce closes a support case in 2026, the marginal cost is different in category, not simply in magnitude. The agent reasons. It calls a tool. It gets a bad answer. It reasons again. It tries another retrieval. It composes a response. Somewhere between ten and a hundred model calls happen. Each one runs on somebody's H100 or B200 cluster. Each one produces output tokens that Salesforce paid for at Anthropic or OpenAI's wholesale rate before it billed you.

That cost lives on Salesforce's cost of goods sold, not on its R&D. It scales with the buyer's usage, not with Salesforce's engineering headcount. If your team runs the agent more, Salesforce's COGS grows. If your team runs it less, Salesforce's COGS shrinks. The vendor's marginal cost is now a linear function of the buyer's behavior in a way that has literally no analog in the last twenty-five years of enterprise software. That single accounting fact is the origin of every pricing debate the industry has spent 2026 arguing about.

The Pricing Debate Is A Symptom, Not A Cause

Every AI vendor spent the last twelve months arguing about how to price. The debate looks like it is about business models. Read it again as an accounting problem.

Salesforce priced Agentforce at two dollars per conversation at launch. Buyers gamed it immediately, because "conversation" is a fuzzy noun. A conversation could branch. A conversation could linger. A conversation could touch three cases and only resolve one. Salesforce introduced Flex Credits, a per-action alternative. Now you buy a bucket of small units and the platform decrements them as the agent updates a record, summarizes a case, drafts an email. Salesforce did this because per-conversation billing did not track per-conversation COGS, and the delta was showing up on Salesforce's income statement.

Intercom went a different way with Fin. Fin charges 99 cents per outcome, where an outcome is a resolved customer support ticket. Standalone Fin ships with a $49 monthly base and 50 included resolutions. Above that, each resolution is metered. Intercom worked directly with Stripe to build the billing infrastructure for this and publicly frames it as outcome-based pricing.

Read the pricing sheet as a hedge. The base fee covers reasoning overhead on inconclusive sessions where no outcome fires but the model still ran. The per-outcome charge covers the successful loops. The buyer pays for the meter. The vendor gets its COGS off its own books. Zendesk uses a similar autonomous resolution unit. Sierra pushes hard on outcome pricing as its foundational thesis. Klarna, when it validated its AI customer service agent against 700 human-equivalent interactions, produced the case study every subsequent outcome-pricing pitch cites for evidence.

The unifying signal is buried in the accounting. Vendors are trying to make revenue track COGS on a per-transaction basis, because the alternative is watching gross margin blow up when a heavy user shows up.

Per-Seat Broke First

Per-seat pricing was the first casualty. If you sell a seat for a hundred dollars a month and one user runs the agent five times a day while another user runs it five hundred, the light user is subsidizing the heavy user's inference bill. That worked in SaaS because the marginal cost of the heavy user's clicks was rounding error. It does not work with agents.

Watch what the LLM vendors themselves did to their own subscription tiers. Anthropic ships Claude Pro at $20 a month, Claude Max at $100 or $200. On August 26, 2026 they published a specific overhaul. Each subscription now comes with a monthly Agent SDK credit pool, and once you burn through it, you pay API rates. That is a public admission that a flat seat could not contain the tail of heavy users. The vendor is telling the buyer, explicitly, that the meter is on.

For enterprise buyers who inherited a decade of assumed per-seat habits, this creates a new procurement problem. Your CFO understands a seat license. She has budgeted seat licenses for years. She does not understand an SDK credit pool with rollover rules and rate-limit interactions. She has never underwritten a variable line where the range between two comparable business units might be ten to one. The number of finance teams that can competently underwrite an agentic AI budget in August 2026 is small.

Cheap Tokens, Expensive Turns

The Gartner note is worth quoting more carefully. Token prices are on track to fall roughly 95 percent by 2030. Per-workflow inference costs for agentic systems will rise more than fivefold over the next two years. Both statements are true simultaneously, and the mechanism is not mysterious.

A chatbot answers a query with one round trip. A reasoning agent answers a query with a plan step, a retrieval step, a tool call, a critique step, a self-correction, a second tool call, and a final composition. Each step is a full model round trip. Some steps involve chain-of-thought traces that push output token counts by an order of magnitude. Multi-step agent workflows run five to twenty-five times the inference of a chatbot for a task the customer perceives as roughly comparable in complexity.

Compounding that, model providers have shipped increasingly capable reasoning tiers at substantially higher per-token prices than their non-reasoning peers. OpenAI's premium reasoning tier prices at $5 input and $30 output per million tokens. Anthropic's Fable 5 sits at $10 input and $50 output. The industry curve is bimodal. Commodity chat tokens are cheap and getting cheaper. Premium reasoning tokens are expensive and holding value. Agents route more of their volume to reasoning, not less, as they take on harder tasks. The mix shifts against the average price.

The result is what Gartner named. Per-unit prices fall, per-task consumption rises faster, and the enterprise inference bill compounds. This is the current run rate at every AI-native vendor with a meaningful book.

The Vendor P&L In August 2026

OpenAI, according to widely-reported figures, spent about $1.35 for every $1 of revenue in 2025. Roughly $3.7 billion in revenue against roughly $5 billion in operating losses. The company sells tokens at a markup over its own compute cost, but the cost of R&D on frontier models, the cost of subsidized capacity for developer growth, and the cost of enterprise seats sold below the wholesale price of the tokens those seats consume all sit above the gross margin line.

Anthropic decided on August 26, 2026 to hold Claude Sonnet 5 at $2 input and $10 output per million tokens indefinitely, rather than raise it to the previously scheduled $3 and $15 on September 1. That decision reads as competitive pressure. It also reads as a signal that the wholesale token market is pricing above the actual cost of running the model, and Anthropic has room to hold the line without operating losses per unit. The middle of the reasoning-model market has margin. The frontier of the reasoning-model market subsidizes the middle. This is the same shape as the traditional cloud market, where compute is the loss leader and value-added services carry the margin, except the loss leader is now a strategic weapon in a five-vendor oligopoly race.

Below the model vendors sit the application vendors. Every AI-native application company is buying tokens at wholesale, running a reasoning loop with a compounding number of round trips, and trying to price the result in a way that covers cost and still looks familiar to a buyer trained on per-seat SaaS. That is why the Bessemer number and the ICONIQ number both put AI gross margins in the 50s. The full-stack economics have compressed by thirty points and there is no scale magic that unwinds it. More users means more inference means more cost. Efficiency in the model layer helps at the margin. The buyer's expanded appetite for agent turns swamps the efficiency.

What The Buyer Actually Owes

If you are a C-level executive evaluating agent vendors right now, the correct mental model is this. You are renting a metered service dressed as software. Every process you deploy an agent against carries a variable cost that will move with your business volume, your team's inventiveness, and the vendor's model routing decisions. You control none of those three completely.

The pricing sheet the vendor puts in front of you is trying to expose that variable cost in some form. Outcome-based pricing exposes it per resolved case. Per-action pricing exposes it per tool call. Per-conversation pricing exposes it per session. Every one of these is a proxy for the underlying inference bill. The proxy is imperfect. When it is imperfect in your favor, the vendor eventually adjusts and closes the gap. When it is imperfect against you, you pay the delta forever.

The buyer's job in August 2026 is to underwrite the proxy. What does an outcome cost you when the agent works? What does it cost when it fails and retries three times before you route to a human? What does the cost distribution look like across your ten most common use cases? What does the tail of your heaviest users look like? If you cannot answer these questions before you sign, you are buying a variable-cost line item with no controls.

I have watched buyers walk into pilots without underwriting the proxy and end up with monthly bills three to five times their pilot estimate. Nobody was fraudulent. The buyer priced the pilot at average usage and hit peak usage in production. The vendor's meter did exactly what the meter said it would do.

The Architecture Question Under All Of This

The pricing debate is downstream of an architectural fact. A reasoning agent is a loop, and every turn of the loop calls a model. The loop's shape is determined by how the agent is designed. A well-designed agent uses cheap models for cheap subtasks, routes to expensive models only when reasoning depth is required, caches aggressively, and terminates loops early when it has sufficient confidence. A poorly-designed agent hits the expensive model every turn, re-fetches context it already had, and runs to loop-cap on tasks it should have exited three turns ago.

The difference in inference cost between a well-designed agent and a poorly-designed one, for the same nominal task, is easily an order of magnitude. That is what any team that has built a production agent has seen in its own token budget. And it maps directly to the buy-versus-build decision every enterprise now faces.

Route Or Be Routed

If you buy an off-the-shelf agent from a vendor, you are inheriting that vendor's routing choices. The vendor is not always incented to optimize for your total cost. If they are on a per-outcome model, they will optimize aggressively to close outcomes cheaply, which sometimes means closing them badly. If they are on a per-token pass-through, they have every incentive to be verbose. Read the pricing structure as a signal of what the vendor's engineers are actually being asked to minimize.

If you build your own agent from scratch, you own the routing, the caching, the termination conditions, the model selection per subtask. You also own the operational burden and the evaluation surface. Most enterprises should not build a general-purpose agent from scratch in 2026. Most enterprises should architect an agent orchestration layer, where their own routing choices sit above a mix of model vendors and specialized third-party agents. The routing layer is where cost discipline lives. The routing layer is where your gross margin lives.

The COGS Line Is Now A Strategy Question

Return to the accounting frame. Cost of goods sold used to be a plumbing question in software. Now it is a strategy question. The gross margin you can sustain determines the R&D you can fund, the sales investment you can support, and the valuation multiple the market will pay you. If your operating model requires 80-point gross margins to work and your infrastructure choices deliver 50, you will be outcompeted by an operator who architected for 65 and is willing to run at 55.

On the buy side, the same math applies. If your unit economics require an internal process to cost X per transaction and your agent vendor charges you 3X, you will be outcompeted by a competitor who architected their process around the actual cost of inference. In an economy where the loop is the work, the loop's economics are the strategy.

The companies that will win the next decade are the ones that treat the reasoning loop as an economic object. They measure its cost. They architect its routing. They price its output to the customer with margin discipline. They rebuild their internal processes to sit inside its actual cost curve, not the cost curve of the pre-agent world.

Model access is not a competitive advantage. Model access is a commodity in a market where five vendors are within striking distance of each other on capability and three of them ship public APIs. The economics of AI agents will be won and lost one architectural decision at a time, above the model layer.

Do Not Lease The Meter

Here is the practical action. Every enterprise leader reading this should be asking one specific question of their AI initiatives. Where does the meter live? Who reads it? Who has the authority to change routing decisions when a subtask starts running expensively?

If the answer is "the vendor decides", you have leased the meter. Your gross margin is a function of decisions made in another company's engineering standup. You will find out about routing changes in your monthly invoice.

If the answer is "we own the routing above the vendor", you have architected the meter. Your gross margin is a function of your own engineering choices, and the vendor is a commodity input you can swap.

The gap between these two states is the single most important decision your AI strategy will make in 2026. It matters more than model selection. It matters more than the specific agent framework. It matters more than the pricing model. Model, framework, and pricing are all variables that either sit inside your control loop or sit outside it. Everything downstream flows from that one architectural fact.

Every quarter this year, another wave of enterprises will discover that their AI vendor's price changed under them, or their usage grew faster than budget, or a new reasoning model cost double what the old one cost per equivalent task. Every quarter, the vendors will adjust their meters, their outcome definitions, their credit pool sizes. The buyer that owns the routing above the meter treats those adjustments as noise. The buyer that leases the meter treats them as strategic threats.

The Case For Architecting, Not Buying

The economics of AI agents will not settle for years. The Gartner projection through 2028 is a headline number. The underlying story is that agent architectures are still being invented, model markets are still being consolidated, and enterprise procurement processes were built for a cost structure that no longer exists. This is the worst possible environment in which to buy a canned solution and hope the vendor holds the line on your behalf.

You need an operator perspective on your side. Someone who reads pricing sheets as accounting statements. Someone who architects the routing layer above the meter, not the workflow that sits under it. Someone who has watched real production agents blow their budgets in real ways and knows how to design orchestration that will not. Someone who treats the reasoning loop as the strategic object it now is.

That is the work Agor AI Advisory does. Not a tool. Not a template. A durable architecture for how your organization will pay for reasoning across the next five years, engineered so the meter belongs to you and the gross margin recovers point by point as you learn.

Sources

Want this kind of automation working for your business?

Agor AI designs and ships the systems these posts describe, scoped in weeks, not quarters.

Book a Free Strategy Call