On September 29, 2026, OpenAI used its DevDay stage to ship two agent products that pull in opposite directions. The first was Dots, persistent agents that, in TechCrunch's report, pursue user-defined goals in the background "with minimal oversight." The second was the Decisions API, an endpoint that writes nothing at all. You hand it a question and a fixed list of allowed answers, and it hands back one of them. Dots got the headlines and the avatar. For anyone responsible for enterprise AI agent deployment, the Decisions API matters more, because it makes a bounded choice almost free and leaves the executive holding a question no model can answer for them. What are the allowed answers?
Most companies have never written that list down. The work of writing it is the strategy now.
Two launches, one afternoon
The details are worth getting straight. According to Unite.ai, OpenAI moved the Decisions API to public beta on October 6, 2026, served through a dedicated POST /v1/decisions endpoint, with gpt-6-luna as the only supported model. The price is $0.10 per million input tokens. Output tokens are free, which tells you how OpenAI thinks about the product. The answer is a few bytes. The cost lives entirely in the context you send.
OpenAI's own speed figures, as reported by NYU Shanghai's research technology group, put a Decisions call at about 150 milliseconds against about 1.6 seconds for a standard Luna call. Nobody has measured that independently yet, and OpenAI has not said what task produced the number. Treat it as a direction rather than a promise.
OpenAI was also second to this idea. TypeSafe AI launched a model called Jev on September 15, 2026, after two years in stealth. Let's Data Science describes it as a model for "structured decisions inside software rather than open-ended text generation." Jev takes typed questions and returns choices, scores, or true-or-false answers, each with a probability, evaluated in parallel rather than token by token. The Register lists its price at $0.042 per million input tokens with output free, and TypeSafe claims 70 to 500 milliseconds end to end. The company says it raised a $40 million seed. It also briefly could not serve its API after launch because demand outran capacity, which is the most honest market signal in the whole story.
Two weeks, two vendors, one shape. A model that picks from a list.
A choice wearing a paragraph
Walk the floor of any operations team and count the decisions. A claims adjuster routes a file to one of six queues. A support agent tags a ticket as billing, technical, or sales. A procurement analyst marks an invoice approve, hold, or reject. A compliance reviewer flags a message as fine or escalate. Most enterprise work that people describe as "judgment" ends in a pick from a short list that already exists, usually in someone's head.
For three years we have pointed text generators at these choices. The pattern goes like this. A prompt asks a large model to read a ticket and "respond with one of the following categories." The model writes a sentence, sometimes a paragraph, sometimes the category with a polite preamble. A parser strips the prose and hunts for the label. A retry loop catches the cases where the model invented a seventh category. An engineer writes a regex for the time it answered in French.
Every one of those steps is a tax on a mismatch. The business needed a choice. The tool produced language. Teams built scaffolding to turn language back into a choice, then wrapped the scaffolding in monitoring because the conversion failed often enough to matter.
The decision models remove the mismatch at the source. The NYU write-up quotes FourWeekMBA's line that the approach "limits the output space itself instead of filtering outputs after the fact." That sentence describes a change in where control sits. In a prompt-and-parse system, control sits in the parser, after the model has already said something. In a decision model, control sits in the list, before the model says anything.
Adam Jacob, CEO of Swamp Club, put the mood bluntly in a post The Register quoted: "It's time to move past the idea that what we need is smarter frontier models." I would put it differently. The frontier models are fine. The bottleneck was never intelligence. It was the absence of a written answer space for the model to be intelligent inside.
The enterprise AI agent deployment problem nobody wrote down
Gartner predicted on June 25, 2025 that more than 40% of agentic AI projects would be canceled by the end of 2027, citing rising costs, unclear business value, and weak risk controls. In the same release Gartner estimated that only about 130 of the thousands of vendors calling themselves agentic were the real thing. Fifteen months later, those three reasons still read like a field report from most of the agent pilots I see.
All three trace back to the same missing artifact. Costs rise because an open-ended agent burns tokens deliberating over choices that should take one pass. Value stays unclear because nobody defined what a correct outcome looks like, so nobody can count correct outcomes. Risk controls stay weak because you cannot fence a space you never drew.
The decision models make the missing artifact visible. To call the endpoint at all, you must type out the options. That act of typing is where most enterprise agent deployment programs quietly fail, long before any model is involved. Ask a claims operation for the complete list of dispositions a pended claim can receive and you will get three versions from three managers, plus a fourth that lives in a 2019 spreadsheet. Ask a bank's onboarding team for every reason an application can be held and someone will say "it depends," which means the list exists in people and nowhere else.
For a decade, that vagueness was cheap. Humans absorbed it. An experienced adjuster knew that "hold for review" and "pending documentation" meant the same thing in practice, and that the regional office used a different word for both. The ambiguity cost a little time and a lot of training. Nobody had to resolve it, because the people doing the work resolved it a thousand times a day without writing anything down.
An agent cannot absorb ambiguity that way. It can only reproduce it at speed. So the first deliverable of a serious enterprise AI agent deployment is a document: a menu of allowed outcomes for each decision the agent will touch, owned by a named person, versioned like code.
The price of a choice fell through the floor
Run the arithmetic on your own volumes, because it changes which projects make sense. Say your support operation routes 2 million tickets a month and each ticket, with its context, runs to 1,000 input tokens. Say you send all of it through OpenAI's Decisions API at the published $0.10 per million input tokens; that is 2 billion tokens and about $200 a month, with nothing billed for the answers. Say you sent the same volume through Jev at $0.042 per million; the bill drops to roughly $84.
Those are rounding errors on any operating budget. The model cost of a bounded decision has gone close enough to zero that it stops being the constraint. Something else has to be the constraint now, and it is the quality of the menu and the quality of the context you feed it.
This flips the usual sequence of an agent business case. Most cases I review start from the model, estimate a per-task cost, then hunt for tasks big enough to justify it. With bounded decisions this cheap, you start from the inventory of decisions, because almost every one of them clears the cost bar. The question becomes which ones you have defined well enough to automate.
Calibration is a business dial
The more interesting feature of Jev is the probability attached to each option. TypeSafe trains for what it calls Reinforcement Learning for Calibrated Decisions, aiming for probabilities that mean what they say: when the model reports higher confidence, it should be right more often. TypeSafe's claims are its own, and Let's Data Science notes they have not been validated across a broad range of production workloads.
If calibration holds even roughly, it turns the old automation debate into a pricing exercise. Suppose a wrong routing decision costs you $40 in rework and a human review costs $6. Suppose the model is calibrated, so a 0.85 confidence means about 15% of those calls are wrong; the expected error cost is about $6, so 0.85 is your break-even line, and anything above it goes straight through while anything below goes to a person. That threshold is a finance decision. It belongs in the same meeting where you set credit limits and refund policies, and nowhere near a prompt file.
OpenAI's version is less clear on this point. The NYU write-up found no confidence or probability output in the Decisions API as of September 30, while The New Stack's coverage said it returns confidence scores. Ask your vendor before you design around it. A decision model without trustworthy confidence gives you speed. A decision model with it gives you a dial.
Reworked's DevDay coverage made the necessary caveat plainly: "The constraint applies to the output, however, and doesn't guarantee the model will choose correctly." A bounded answer can still be the wrong answer. What the bound buys you is that every wrong answer is a known kind of wrong, which you can count, price, and route.
Who owns the menu
Once the menu is the control surface, the person who writes it holds real power, and most org charts have no seat for that person.
In the prompt era, the person closest to the model's behavior was usually an engineer, because the behavior lived in prompts and parsers. In the menu era, behavior lives in the list of allowed outcomes and the thresholds attached to them. Adding a disposition called "refer to fraud" changes how the company treats customers. Removing "manual override" changes who has discretion. Merging two refund categories changes what finance can report. Those are policy decisions dressed as configuration.
So the owner of each menu should be whoever owns the policy it encodes. The head of claims owns the claims dispositions. The controller owns the invoice outcomes. The general counsel's office signs off on any menu that touches regulated communication. Engineering maintains the plumbing and keeps the menus in version control, the way it keeps infrastructure definitions there.
That sounds like bureaucracy. In practice it removes bureaucracy, because it replaces a dozen informal understandings with one file that changes through review. When an auditor asks why the system did something, the answer is a diff with a name and a date on it.
"None of the above" is the most important option
Every menu needs an exit. Call it escalate, unknown, or out of scope, but it must be on the list, and it must be cheap for the model to choose.
A menu without an exit forces the model to pick the least bad answer for cases that fit nowhere. That is how bounded systems produce confident nonsense. The exit option is where the world tells you your menu is wrong. Watch its volume. When the share of cases landing in "none of the above" climbs, something new is happening in your business that your categories do not describe, a new fraud pattern, a new product, a new kind of customer complaint. In most companies that signal currently reaches leadership months late, filtered through anecdotes. A well-built menu delivers it as a weekly number.
Let's Data Science's advice on Jev lands in the same place: labeled holdout data, escalation thresholds for uncertain answers, and drift monitoring before automated decisions go live. Those three items are cheap. Skipping them is how a bounded system ends up worse than the unbounded one it replaced.
The menu goes stale
There is a cost to all this, and executives should see it clearly before signing up.
A menu freezes a version of your business. The categories you write today encode today's products, today's regulations, today's org chart. Humans updated those categories informally, by drifting. A clerk started using "hold" for a new situation and the practice spread before anyone named it. Bounded agents do not drift. They keep choosing from the list you gave them, with perfect consistency, long after the list has stopped fitting.
That consistency is the reason to use them and the reason they age badly. The remedy is a maintenance cadence. Every menu gets a review date. Every review looks at three numbers: exit-option volume, override rate by humans downstream, and the distribution of choices over time. A category that nobody picks for a quarter is dead and should go. A category that absorbs a growing share of cases is probably two categories now.
Companies already run this discipline for chart-of-accounts codes and product SKUs. Decision menus belong in the same family. They are master data, and they deserve a master-data owner.
Where the open-ended agent still belongs
Dots and decision models sit at opposite ends of a spectrum, and OpenAI shipping both on the same day is a fair map of where enterprise AI agent deployment is heading.
Open-ended agents remain the right tool for work whose outcome space is large or unknown. Drafting a contract clause, investigating an incident, assembling a board pack, exploring a dataset nobody has looked at. In those tasks, writing a menu in advance would mean writing the answer in advance. Reworked notes that for Enterprise, Edu and Healthcare customers the Dots beta is off by default and has to be switched on by a workspace admin, which is a sensible posture for agents whose action space is wide open.
The architecture I recommend to clients puts bounded decisions at every junction and open-ended reasoning between them. An open-ended agent can investigate a disputed invoice for as long as it needs. When it reaches a point of consequence, such as paying, holding, or escalating, it hands off to a decision call against the controller's menu. The open-ended part does the thinking. The bounded part does the committing. TypeSafe's own team suggested a version of this, according to The Register: using Jev to route tool and MCP calls for other models.
That split also answers the governance question that has stalled so many agent programs. Boards worry about agents taking actions nobody approved. With this design, every consequential action passes through a list that a named executive signed. The agent can be as creative as it likes in reaching the junction. It cannot invent a new exit from it.
What changes on Monday
The shift here is small enough to miss and large enough to reorder your agent roadmap. For three years, the hard part of putting AI into operations was getting a model to behave. As of October 6, 2026, a model that picks from your list costs about $0.10 per million tokens of context and answers in a fraction of a second, if OpenAI's figures hold. The hard part has moved upstream into your own operating definitions, where it was hiding the whole time.
Start with an inventory. List every recurring decision in one function, the ones that end in a pick from a set. Most functions have between a few dozen and a few hundred. For each one, write the menu, name an owner, add the exit option, and attach a cost of error. Then sort by volume times error cost. The top of that list is your first deployment, and it will almost never be the flashy autonomous agent someone demoed at the offsite.
Then build the junction architecture before you build agents. Decide where open-ended reasoning is allowed and where it must commit through a menu. Put the menus in version control. Give every threshold a finance owner. Put exit-option volume on an operating dashboard where executives will see it every week.
None of this is a purchase. Jev and the Decisions API are commodities already, two vendors priced within cents of each other within three weeks of the first launch, and a third will follow. What does not commoditize is the menu itself: your company's written, owned, current statement of what it is allowed to decide and how sure it must be before it does. Companies that write that document will run cheap, fast, auditable agents across every function. Companies that skip it will keep paying frontier-model prices to have a paragraph written and then parsed back into a choice they could have specified on day one.
That is an architecture problem, and it touches policy, finance, compliance, and engineering at once, which is exactly why it falls between the chairs of most leadership teams. Agor AI Advisory works with executive teams to run the decision inventory, design the junction architecture, and set up the ownership and review cadence that keeps menus alive after launch. Your competitors bought the same endpoints you can buy this afternoon. The menu is the part they cannot copy.
Sources
- Unite.ai, OpenAI Decisions API public beta, October 2026
- NYU Shanghai RITS, OpenAI's Decisions API Picks Answers Instead of Writing Them, September 2026
- TechCrunch, OpenAI launches Dots, its bubbly agentic avatar, September 29, 2026
- Reworked, The Biggest Announcements From OpenAI DevDay, 2026
- The Register, Shut up and calculate: Jev's new AI primitives for coders, September 23, 2026
- Let's Data Science, TypeSafe AI Launches Jev for Structured Software Decisions, September 19, 2026
- Gartner, Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, June 25, 2025
