Between January 23 and October 3, 2026, I ran Claude Code on my main machine across 254 days, 46,519 sessions and 905,244 messages. It processed 89.0 billion tokens. Over the same months I paid Anthropic $2,019.92: one $20 Pro renewal and eight renewals of the Max 20x plan at $249.99 each, billed through Apple.
This weekend I priced every one of those tokens at Anthropic's published API rates, model by model, the way a company building on the API would be billed. The total is $90,245. That is 45 times what I paid. Counting every Claude subscription I have bought since March 2024, $2,299.92 in all, it is still 39 times.
I ran the same exercise on HeyGen, the AI video service I use for avatar video. Its 944 finished videos (389 minutes) and 108 avatars come to between $2,211 and $3,863 at HeyGen's API rates, against $271 paid since early 2025: 8 to 14 times. Then I went through every AI receipt I have, back to May 2023: twelve services, $4,721.35 in all. Together they delivered between $92,931 and $95,015 of API-priced work, about 20 times. Almost all of it came from one of them.
I am publishing my own numbers because they are the clearest way I know to answer a question clients ask me every week: should we buy seats, or pay for the API? The answer depends on where the value concentrates. In my data it concentrates in three places that almost nobody looks at.
How I counted
Claude Code keeps its own usage statistics, the cache behind its /stats command. For every model it records lifetime input, output, cache-write and cache-read tokens, and since August 7 it has also recorded each model's tokens per day. I multiplied each count by that model's list price per million tokens. Cache writes bill at 1.25 times the input price for a five-minute cache and twice the input price for a one-hour cache, per Anthropic's prompt caching documentation. 77% of my cache writes used the hour, a share I measured from the session transcripts still on disk. Cache reads bill at a tenth of the input price or less.
The true figure is higher than this, not lower. It leaves out every conversation in the Claude apps on the web, desktop and phone, everything on my other machines, and everything before January 23, when the statistics begin. One thing makes the monthly picture less precise than the total. Before August 7, Claude Code kept only lifetime totals per model and a count of messages per day, so I spread each older model's tokens across the days it was in use, in proportion to the messages I sent. The lifetime total for each model is exact. Which month a given dollar landed in, before August, is an estimate.
Most of the value is cache
Of the 89.0 billion tokens, 83.8 billion (94.2%) were cache reads: the model rereading context it had already seen, at a steep discount. Cache reads came to $37,597, 42% of the bill. Cache writes, 4.5 billion tokens, came to $41,272, or 46%. Output, the text and code the model actually wrote, was 380.7 million tokens, 0.43% of the total, and $9,666, 11% of the bill. Uncached input was $1,710.
That is what agentic work looks like. A coding agent does not answer a question and stop. It reads the repository, runs a command, rereads everything with the result attached, and goes again, sometimes hundreds of times in one task. Every turn carries the whole conversation forward. The output is small. The context is enormous, and it repeats.
It also means caching is the price. Billed as plain input, those 83.8 billion cache reads alone would have cost $425,630 instead of $37,597. With no caching at all, the same nine months would have come to about $460,000, five times the cached bill. A company running agents on the API that does not design for cache hits (a stable system prompt, context that only ever grows at the end, tool definitions that stay fixed between turns) is not paying 45 times a subscription. It is paying five times that.
Most of the value is a few months
The second surprise is how uneven it was. June ($23,219) and September ($23,083) together were 51% of the total. In those two months an average day carried about $766 and $761 of value above what the plan cost per day. In August the figure was $158. In March, my first full month on Max, I used about $805 of API-priced work against a $256 share of the fee, roughly three times what I paid. The busiest 25 days, a tenth of the period, carried 44% of the value, and on 81 days I barely used it at all. The biggest single day, June 11, came to $2,335, more than nine months of Max.
A flat rate is priced on an average user. My usage was not average in any month, and it swung by a factor of nearly 30 between the lightest full month and the heaviest. The plan absorbed both. On the API, June and September are the months a finance team calls a meeting about.
Most of the value is the frontier
84% of the value ran on Opus models and 13% on Fable, Anthropic's highest-priced tier. Sonnet and Haiku together were under 3%. Opus 4.8 alone was 31.8%, Opus 5 another 19.4%. On a flat rate there is no price signal telling me to step down to a cheaper model, so I mostly did not. On the API, every one of those calls is a choice somebody has to justify, and most teams would send a good share of them to a smaller model.
What 45 times does not mean
It does not mean I saved $88,000. At API prices I would not have done the same work. I would have run fewer agents in parallel, kept contexts shorter, sent more tasks to Sonnet, and stopped some experiments sooner. Some of that would have been discipline and some of it would have been worse work. The $90,245 measures how much computing a flat rate let me use without thinking about it. That is a different thing from money in my pocket.
It does not mean the deal is permanent, either. Max plans come with usage limits that Anthropic sets and can change, and the plan is priced for a population. That population includes plenty of people who use a fraction of what they pay for. Their quiet months pay for my Junes. A subscription this lopsided is a bet the vendor is making on the average, and vendors revise their bets.
And it is one person's data: a heavy user, a coding workload, one vendor's price list. A sales team drafting emails in a chat window would see a far smaller multiple, quite possibly below one.
HeyGen is the useful contrast. Its multiple is 8 to 14, not 45, and 704 of the 1,653 renders on the account (43%) failed and are not counted at all. Avatar video has none of the cached, endlessly reread context that makes coding agents so cheap per unit of value, and a failed render costs time even when it costs no credits. Flat-rate video is still a good deal for me. It is an ordinary good deal, not an extraordinary one.
Every other AI service
Claude and HeyGen are the two services whose usage records were good enough to price line by line. They are not the only AI I pay for. So I pulled every AI charge I could find in Apple purchase history, card statements and each API's billing page, and priced the usage wherever a service keeps a record of it: ElevenLabs from its usage dashboard, OpenAI from its costs API, Gemini from Google Cloud's request metrics and invoices, and xAI rebuilt from my own narration files and call logs.
Claude was 49% of what I spent and 97% of the value. HeyGen is the only other service clearly above one times. The pay-per-use APIs (ElevenLabs, Gemini, OpenAI, xAI) all land between about 0.4 and 2 times, which is what pay-per-use should look like: you pay roughly list price for what you use, and the gap below one is prepaid credit or a monthly allowance that went unspent. OpenAI shows it plainly. I bought $125 of credit and used $99.04 of it. ElevenLabs shows it at scale: $376 across its monthly plans for $158 to $311 of voice at its API rates.
The last 29% of the spend, $1,382, went to consumer apps: ChatGPT, the Gemini app, Midjourney, Suno, Runway and SlideSpeak. None of them keeps a usage record I can price, so here they count as paid with no value. Some were worth it. I cannot show which, and that is the point. For nearly a third of my AI spend I had no way to answer the question this post asks.
One more thing turned up. The Gemini numbers looked odd: 6.05 million tokens and 250 requests to a single model, all on October 3. The request pattern (streaming calls, preset lookups and token counts from Google AI Studio's own client) says that was me building an app in AI Studio, billed to the API key it was attached to. A prototyping tool spending real API money is easy to miss on an invoice.
The cap, not the cash
Measured this way, my cash subscriptions are already lean. ElevenLabs is back on its free tier, HeyGen lapsed in August and ChatGPT stopped in July. What actually limits me is the Max plan's weekly usage cap. I have hit it three weeks running, and in ten days the scheduler that runs my unattended jobs skipped 125 of them to stay under it.
So the useful question is not which subscription to cancel. It is which work deserves the capacity that returns 39 times its price. I went through all 51 places my systems call a model and found twelve changes, mostly moving bulk research, scoring and extraction off Claude to Gemini Flash, with Google Search grounding or batch pricing where they fit. Together they free about $152 a month of Max capacity at API prices, and Gemini Flash does that kind of work for cents a job. Judgment, planning and anything written in my voice stay on the frontier.
What this means for buying AI
The practical question for a business is seats or API. My numbers point to a simple rule and a few corollaries.
Seats for people, the API for processes. A subscription is the cheapest way to give a heavy individual user frontier models, because the vendor averages your power users against everyone else's light ones. An engineer or analyst who lives in an agent all day is exactly the user a flat rate is generous to. The API is for work that runs without a person at the keyboard: pipelines, customer-facing agents, anything on a schedule or a webhook. Seats are built for an individual at a keyboard, not for a service.
Measure before you renew. Most teams buy seats on headcount and never look again. Price your heaviest users' actual usage at API rates for one month. Claude Code's cost documentation shows where to find the numbers. If a seat's API-equivalent is below its price, that seat is funding the vendor's averages. If it is ten or forty times the price, you have found the people whose work the subscription is quietly carrying, and you should know who they are before the limits change.
On the API, design for the cache first. In my data, cache traffic was 87% of the cost, and without caching the whole bill would have been five times larger. Prompt structure is a cost decision, not a style choice.
Route by task. Almost all of my value ran on the most capable models. On a flat rate that costs nothing extra. On the API it is the second biggest lever after caching. Send planning and hard reasoning to the frontier and the mechanical steps to smaller models. My own unattended jobs already work that way, on three tiers by role.
Make every AI line item measurable. Nearly a third of my spend sat in apps with no usage record. If a tool cannot tell you what you used, you cannot tell whether it pays for itself. Prefer tools that can, and review the ones that cannot each renewal.
When a flat rate is the cheap thing, spend it on the hard work. A subscription that returns 39 times its price becomes a capacity problem, not a cost problem. Route mechanical work to a cheap pay-per-use model and keep the flat-rate capacity for the work only the frontier can do.
Budget for spikes. Two months were half of my year. An API budget set on an average month would have been blown twice. Either cap it, or put spiky, exploratory work on seats and steady production work on the API.
Agor AI Advisory helps companies make exactly this call: which people should have seats, which processes belong on the API, how to measure the difference with your own usage, and how to build agents whose caching and model routing keep the bill where you planned it. Schedule a strategic consultation with us today.
Sources
- Anthropic, Claude API pricing
- Anthropic, Prompt caching documentation
- Anthropic, Claude plans and pricing
- Anthropic Help Center, "What is the Max plan?"
- Claude Code documentation, "Manage costs effectively"
- HeyGen API pricing
- ElevenLabs API pricing
- Google, Gemini API pricing
- OpenAI API pricing
- xAI API models and pricing