On September 22, 2026, Anthropic released Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens. According to VentureBeat's launch coverage, that is 20% below Opus 5 and 60% below Fable 5.1, the model it beats on Terminal-Bench 4.0 (66.4% to 55.8%). Anthropic also says typical workloads come out about 40% cheaper once you count the model's better token efficiency.
Five days earlier, on September 17, Nebius told customers that on-demand GPU prices would go up on October 1. Investing.com reported the notice covered Nvidia H100, H200, B200 and B300 capacity, plus some CPU and memory resources. The Motley Fool added up the figures on September 27: an H100 hour goes from $3.85 to $4.50, and a B300 hour from $7.85 to $9.50. It is Nebius's second increase since May, when an H100 rented for $2.95. That comes to 53% in five months.
So in one September week the price of thinking fell and the price of the machines that think went up. Most finance teams I talk to have one line for AI, and it sits under software. That line assumes a fixed annual price, set once and paid in twelve equal slices. I think the assumption is now wrong, and it is wrong in a way the CFO already knows how to handle. AI total cost of ownership has started to act like a commodity input. Commodity inputs get a treasury desk.
Two prices, one direction of travel
The two September moves look like they contradict each other. They fit together once you see where each company sits.
Anthropic sells finished tokens. Every model release lets it push more useful work through the same silicon, so it can cut the list price and still protect its margin. Better kernels, better caching and smaller active parameter counts all fall on the seller's side of the ledger. The customer gets the savings as a lower sticker.
Nebius sells raw capacity, and raw capacity is short. Nebius management said, as the Fool quoted them, "we sold out of capacity because, as fast as we bring capacity online, we can sell it." The CEO said the company could sell all of its 2027 capacity today on current terms. Its remaining performance obligations stood at $37.5 billion at the end of the second quarter, more than 60 times quarterly revenue. When a seller can say that, the price goes up.
Put the two together and a buyer faces a strange market. The unit price of finished intelligence is falling fast. The price of the scarce input under it is rising. Whether your own bill goes up or down depends on which side of that split you buy from, how much you consume, and whether you locked anything in. Those are commodity questions, and a software budget has no place to record them.
What AI total cost of ownership actually contains now
The standard AI TCO model came from enterprise software, and it has three parts. You pay to build or buy, you pay to run, you pay to maintain. Consultancies still publish versions of it. The usual rule says maintenance runs 15 to 25% of build cost each year, and three-year TCO lands at 1.5 to 2 times the initial build. Those numbers are fine as far as they go. They describe a system whose running cost is flat.
Running cost stopped being flat in 2026, for two reasons that September made plain.
The seat turned into a meter
GitHub moved every Copilot plan to usage-based billing on June 1, 2026. In the April 27 announcement, GitHub's chief product officer Mario Rodriguez gave the reason plainly: "a quick chat question and a multi-hour autonomous coding session can cost the user the same amount." The flat seat had become a subsidy for the heaviest users, and GitHub stopped paying it.
The new structure keeps the seat fee but turns it into a credit allowance. Copilot Business at $19 per user per month now includes $19 of AI Credits, where one credit is one cent, spent at each model's published token rates. Code completions stay free. Everything agentic draws down the balance.
GitHub cushioned the switch. Business customers got $30 of monthly credits in June, July and August, and Enterprise customers got $70 instead of $39. The cushion ended on August 31. September 2026 is the first month a Copilot Enterprise customer lives on the real allowance, which is 44% smaller than the one they had over the summer. No price changed on September 1, and yet thousands of engineering budgets got tighter that morning.
Nobody announced a price hike here. What changed is the kind of cost. A seat is a fixed cost you can count on a headcount plan. A meter is a variable cost that moves with behavior, and the behavior belongs to your employees and their agents.
Volume is set by the people who don't sign the invoice
Uber is the case everyone now cites, and it deserves the attention. Fortune reported on May 26, 2026 that Uber burned through its entire 2026 AI budget in four months. Uber had encouraged adoption with internal leaderboards that ranked teams by AI usage. It worked. Claude Code use climbed, R&D spending hit $951 million in the first quarter, and chief operating officer Andrew Macdonald was left saying of the link between spending and customer value, "That link is not there yet."
Uber did nothing unusual. It set a fixed annual number for an input whose volume was chosen by 5,000 engineers, then paid those engineers, through the leaderboard, to raise the volume. A treasury team would recognize that at once. The company was short a commodity with no cap on exposure, and it had put a bonus on buying more.
Cheaper tokens will raise your bill
The 20% cut on Opus 5.5 will be read in many boardrooms as good news for the budget. For most of them it will be the opposite.
When a unit of work gets cheaper, people find more units to run. Economists call this the Jevons effect, and inside a firm it moves faster than in a national economy, because nobody has to build a coal plant to act on it. An engineer who ran one agent overnight runs four. A support team that summarized 10% of tickets summarizes all of them. A product manager who never touched the API writes a scheduled job that calls it 20,000 times a day. Each choice looks sensible at the new price, and each one adds a permanent line to next quarter's usage.
Anthropic's claim that Opus 5.5 does typical work for about 40% less makes this sharper. Agent loops are where most token volume now lives, and those are the loops that get cheaper per task. They are also the loops with no natural stopping point. A human asks one question and reads the answer. An agent keeps going until it finishes or hits a limit someone set. If nobody set one, a cheaper model mostly buys you a longer loop.
So a falling unit price and a rising total bill are the normal result of a price cut on an input with elastic demand. Any honest AI TCO model has to predict it. Most don't, because they multiply today's price by today's volume and call it a forecast.
The desk your CFO already runs
Companies that buy jet fuel, copper, natural gas or foreign currency do not put those inputs in a software budget. They run them through treasury, where a small team does four jobs. It maps the exposure, meaning how much the company buys, from whom, and at what terms. It decides how much of that exposure to lock in and how much to leave floating. It keeps the option to switch suppliers or substitutes. And it controls volume from the demand side, so the business cannot run up an open position by accident.
Every one of those jobs now has an AI version, and in September 2026 the vendors started selling the instruments for them.
Map the exposure
Start with a plain count of where your money goes. Separate the four layers. Seats with credit allowances, like Copilot. Direct API spend, with Anthropic, OpenAI or Google. Hosted inference on open-weight models, through providers like Fireworks or Together AI. Raw GPU hours, from Nebius, CoreWeave or a hyperscaler. Each layer has its own price trend, and September showed they can move in opposite directions within days.
Most companies cannot produce this map today. API keys sit on corporate cards. Seat plans sit with IT. GPU contracts sit with a data science team that signed them in 2025. The first job of the desk is to get all four onto one page with monthly volume and unit price, so you can see which way the whole position leans.
Decide what to lock in
Once you can see exposure, you can hedge it, and the market now sells hedges. On September 7, 2026, Fireworks AI launched Reserved Throughput, a dollar-per-minute capacity reservation on its serverless inference with an SLA up to the reserved level. Overage bills at the normal serverless rate with no SLA, and unused capacity does not roll over. Fireworks pitches it at customers who keep hitting rate limits.
A capacity reservation of that kind works like a forward contract. You pay for certainty of supply. You give up some flexibility, since unused minutes vanish. Whether it is a good trade depends on the same things that decide whether an airline should buy fuel forward: how stable your base load is, how much a shortage would cost you, and which way you think prices are heading.
The Nebius hike tells you which way raw capacity is heading, at least for now. A company that signed a year of reserved H100 capacity at May prices has a contract worth 53% more than it was five months ago. A company that stayed on demand pays the new rate on October 1. For token-denominated spend the sign flips. Anyone who locked a long commitment at Opus 5 rates in August now pays 20% more than a new customer for a weaker model. Deciding what to lock is a real call, and it needs someone whose job is to make it.
Keep the switch open
A treasury desk values the option to change supplier because that option caps what any one supplier can charge. For AI, the option is architectural. If your prompts, tools and evaluation suites only work with one model family, you have no option, and every price change passes straight through.
September offered plenty of chances to switch. Google launched Gemini 3.8 Flash on September 4 at $0.75 in and $3.75 out per million tokens, according to the usage-pricing tracker. Cohere released North Mini Code free under Apache 2.0 on September 11. Opus 5.5 landed on September 22 at a lower price than its predecessor. A company with a model router and a regression suite could move a workload to whichever of these fits best within a week. A company without one pays whatever its default vendor charges.
The same tracker shows why the option also protects you from the downside. Moonshot AI retired Kimi K2.5 on August 31, and calls started returning 404 errors. LMNT, a text-to-speech provider, shut down entirely on September 1. A supplier that disappears is the extreme case of a price change, and the only hedge is the ability to move.
Control volume at the source
The last job is the one Uber skipped. Treasury does not let a plant manager commit the company to unlimited copper purchases. In AI, every engineer with an agent and every business user with an API key can commit the company to unlimited token purchases, and most can do it without a single approval.
Volume control does not mean rationing. It means putting spend limits on agent loops, setting per-team monthly allowances that roll up to a number finance can see, and making the cost of a workflow visible to the person who built it. GitHub already did a version of this for you when it turned seats into credit pools. The companies that handle September well will have extended the same idea to every other meter they run.
Whatever you do, stop rewarding raw usage. A leaderboard that ranks teams by tokens consumed is a trading desk that pays traders for the size of their positions and never checks the profit. Rank teams by work shipped per dollar of compute, or by nothing.
The budget cycle is the wrong clock
One structural problem sits under all of this. Most companies set the AI line once a year, in the autumn budget round, for the following calendar year. Look at what happened between one budget round and the next. In May an H100 hour cost $2.95. By October it costs $4.50. Anthropic's flagship price fell 20% in one release. GitHub changed the unit of billing for a product used by millions of developers. Any of these alone would break a fixed annual number.
A commodity input needs a rolling forecast, reviewed monthly, with a range instead of a point. Airlines publish fuel sensitivity in their guidance: this much margin per dollar of movement in the price of a barrel. Finance teams should be able to say the same thing about AI. What happens to operating margin if token volume doubles, if GPU rent rises another 20%, if the default model is deprecated with 30 days' notice? If nobody in the building can answer, nobody in the building owns the exposure.
This also changes how you read vendor announcements. A price cut from Anthropic or Google is information about their cost curve and about the volume they expect you to add. A price rise from Nebius is information about how much capacity is left. The desk reads both the way a fuel buyer reads OPEC statements, as signals about where the position should be next month.
Where most companies will get this wrong
The common mistake will be to treat the answer as a tool purchase. There are plenty of AI cost dashboards for sale, and some are good. A dashboard shows you the bill. It does not decide what to lock in, it does not build the routing layer that lets you switch models, and it does not change the incentives that drive volume. Those are design decisions about how your company buys and uses compute. A vendor cannot make them for you, because the vendor is one of the counterparties.
The second mistake will be to overcorrect. After the Uber story, some companies froze AI spend or capped seats hard. That trades a cost problem for a capability problem, and the capability problem is worse. Anthropic's 40% efficiency claim and Opus 5.5's benchmark gains mean the work per dollar is improving fast. A company that freezes spend falls behind rivals who are buying more capability at lower unit prices. What treasury does is manage the position so you can keep buying without being surprised.
The third mistake is structural, and I see it most often. Companies let AI costs land wherever the purchase happened: engineering, marketing, support, IT. Each team optimizes its own slice, and nobody sees the whole position. Four teams each sign a separate annual commitment with the same vendor at different rates. Two teams run the same workload on different models without knowing it. Meanwhile the GPU contract a data science lead signed in 2025 quietly becomes the best deal in the company, and no one else knows it exists.
Build the desk before the next invoice
The window for doing this calmly is short. Nebius's new rates start on October 1. The Copilot cushion is already gone. Opus 5.5 is live, and your engineers have almost certainly switched to it, which means your token volume is climbing at the new lower price right now. The October invoices will be the first to show all of September's changes at once.
The work is architecture. Someone has to map the four layers of exposure onto one page. Someone has to build the routing and evaluation layer that turns vendor choice into a real option. Someone has to set the allowances and loop limits that put volume under control, and someone has to design the incentives so teams are paid for output instead of consumption. Finally, someone has to set up the monthly forecast that gives the CFO a sensitivity number to report. None of this comes in a box. It changes who decides what, and it has to be built around how your company actually works.
Agor AI Advisory builds this function with executive teams. We have done the exposure mapping, the routing architecture and the governance design across enough companies to know where the hidden positions sit and which hedges are worth their price. The companies that set up a compute treasury this quarter will read every vendor price change in 2027 as a signal to act on, while the rest will find out what happened when the invoice arrives.
Sources
- VentureBeat, Anthropic releases Claude Opus 5.5, September 22, 2026
- The Motley Fool, Nebius Is Raising the Price of Its AI Compute on Oct. 1, September 27, 2026
- Investing.com, Nebius to increase prices for Nvidia GPU resources from October, September 17, 2026
- The GitHub Blog, GitHub Copilot is moving to usage-based billing, April 27, 2026
- Fortune, Uber burned through its entire 2026 AI budget in four months, May 26, 2026
- Usage Pricing, Fireworks launches Reserved Throughput, September 7, 2026
- Usage Pricing, AI pricing changes and updates tracker, September 2026