← Back to Insights

Insight

Ask For the Odds

Ariel Agor
•
Ask For the Odds

Listen · Read by Leo · click any word to jump

0:00 / —· loading…

On September 15, 2026, a San Francisco startup called TypeSafe AI came out of two years in stealth with a model named Jev, after the economist William Stanley Jevons. Jev cannot write you an email. You hand it a situation, a question and a closed list of allowed answers, and it hands back probabilities. Within 24 hours of arriving on Vercel's AI Gateway, roughly 13% of paid teams had used it, double any previous model launch on that gateway. TypeSafe priced it at $0.042 per million input tokens and nothing at all for output.

Two weeks later the copies came all at once. OpenAI showed a Decisions API at DevDay on September 29. On October 1, Cloudflare released Clef and Clef-flash, Perplexity released pplx-decider-v1-27b, and AWS's Strands Labs released Strands Decider 2B. TechCrunch's headline that day said decision models were flooding the web, and the headline understated it.

The adoption curve is the part executives should study. One in eight paying teams on a developer gateway did not discover a new need overnight. They already had a problem. Somewhere in their software, a chat model sat at a fork in a workflow, was asked to pick one of five options, and answered in paragraphs. Jev was a relief valve for a pressure that had been building for three years.

That pressure comes from one of the most common AI implementation pitfalls of this cycle, and the decision-model wave has finally made it visible. Companies put a writer in charge of a switch. Then they skipped the hard work the switch required, which was writing down the decision.

The writer at the switch

Most enterprise AI built since 2023 has a shape that will look odd in five years. A support ticket arrives. Software sends it to a frontier model with a prompt: read this, decide whether it is billing, technical or cancellation, and say how urgent it is. The model writes back a few sentences. A parser fishes the category out of the prose. A developer adds "RESPOND ONLY IN JSON" in capitals, and it mostly works.

Every part of that loop costs money that nobody budgeted for as a design choice. You pay for output tokens that narrate a choice nobody reads. You wait while the model composes. You get a different phrasing each run, so the parser breaks on the edge cases. And you never learn how sure the model was. "This appears to be a billing issue" reads identically whether the model was 98% confident or flipping a coin.

The numbers from the past three weeks put a price on the habit. Cloudflare reported that its own Threat Intelligence team classified domains in 4.7 seconds using gpt-oss-120b and 2.2 seconds with Clef, while getting more classifications back per call. In Cloudflare's internal suite, Clef ran at a median of 209 milliseconds, Clef-flash at 38.8, and Jev at 524. TechTarget set Jev's pricing beside Claude's at $10 and $50 per million input and output tokens. A demo cited by TechCrunch put the cost of monitoring an agent workload at $2.94 with Jev against $372 with a frontier model.

Those gaps are two orders of magnitude wide. Nobody gets to that kind of waste through bad luck. A whole category of software was built on a category error, and the fluency of the models hid it. A chat model will always give you an answer that sounds reasonable. It took a model that refuses to talk to show how much of the answer was talk.

Common AI implementation pitfalls the decision wave exposed

It would be easy to read the last three weeks as a procurement story: swap the expensive model for the cheap one at every fork, pocket the savings. Teams that do only that will fall into a new set of holes, because the decision model does something the chat model never did. It makes you show your work before it will show its own.

Nobody wrote the question down

Every one of these APIs demands the same three inputs. Clef's documentation names the question types plainly: a "noul" asks whether something is true and returns a probability between 0 and 1, a "choice" asks which of several options applies, and a "score" asks where something sits on a scale. Perplexity's model is trained to output a probability distribution over a fixed set of answers. You cannot call any of them without stating, in advance, the exact question and every answer you are willing to accept.

That requirement is where most projects will stall, and it should be. Chat models let teams avoid it. You could send a ticket and the instruction "figure out what to do with this," and the model would produce something plausible. The ambiguity in your operations got absorbed into the model's prose, where nobody had to look at it.

Enumerating the answers is a policy act. Take a refund workflow. Is "partial credit" a separate answer from "refund"? Is "escalate to a human" an answer, and at what point? Is there a "none of these" option, and where does it route? Those are calls about what the company is willing to do for a customer. A product manager with a prompt file has been making them by accident for two years. When a decision model forces the list onto a page, the list belongs in front of whoever owns the P&L for that workflow.

My experience across client engagements is that most companies cannot produce this list for their top twenty automated decisions. They know which workflows have AI in them. They do not know which decisions those workflows make, what the legal options were, or who approved the set.

The threshold has no owner

Say the model returns 0.71 on "is this transaction fraudulent." What happens next?

Somebody has to pick a cutoff. Set it at 0.5 and you block more fraud and anger more honest customers. Set it at 0.9 and you wave through losses to protect conversion. That number is a trade between the cost of a false positive and the cost of a false negative, and it is one of the most consequential settings in the business. With a chat model, the threshold was buried in the model's own unstated sense of when to say "fraudulent." With a decision model, it is a number in a config file, and someone typed it.

The pitfall is letting the person who typed it be whoever happened to be writing the integration. A threshold is a pricing decision in disguise. It needs an owner with a name, a review date, and a dashboard showing what the false positive and false negative rates are actually costing.

Calibration makes this harder than it looks. Cloudflare says it trained Clef with a method it calls Reinforcement Learning for Calibrated Decisions, using Brier loss so that a 0.7 should be right about 70% of the time. That work matters. It also describes Cloudflare's training distribution, not your Tuesday-morning ticket queue after a pricing change. TechTarget's coverage flagged the other caveats: Jev gives no visible reasoning, and its answers can vary across runs. A probability that was calibrated in September drifts by December, and the only way you will know is if someone is checking predicted odds against outcomes. Most companies have no one assigned to that job, because with prose there was nothing to check.

The paragraph felt like an explanation

The quiet comfort of a chat model at a decision point was that it explained itself. Compliance teams liked the paragraph. Auditors liked the paragraph. The paragraph was often a story the model wrote after the fact, fluent and unfalsifiable, yet it felt like accountability.

Decision models remove the story. For a support router, nobody will miss it. For credit, insurance, hiring, or anything under the EU AI Act's high-risk categories, the loss will surface in the first review. The record you need now is plainer and more honest: the inputs, the version of the question and answer set, the probability returned, the threshold in force, and who set it. Companies that leaned on the paragraph for comfort will have to build that record. The ones that never wrote down the question will discover they cannot.

Marrying a model in the middle of a clone war

Marc Brooker, a distinguished engineer at AWS, wrote on September 28 that he had set out to build the best 2B-parameter decision model and succeeded "for about 24 hours before being overtaken in the benchmarks." AWS cleaned up his prototype and shipped it as Strands Decider 2B. TypeSafe's CEO, Diogo Almeida, a former OpenAI researcher, told TechCrunch the new batch "seems more like ML people wanting to implement a cool architecture than a team deeply dedicated to making intelligence useful."

Both men are describing a market about to commoditize in weeks. Clef and pplx-decider ship under Apache 2.0. Cloudflare advertises Clef as fully compatible with Jev's API. When the best model at a task changes every day and the leaders promise drop-in compatibility, the model is the least defensible part of your stack.

So watch what the vendors are really selling. Cloudflare launched Clef alongside a reinforcement learning fine-tuning service. In its own description, AI Gateway captures your request and response data, Workers AI generates rollouts, a sandbox scores them, and a trainer updates the weights, with a forward-deployed engineering team doing it by hand for now. That pipeline is sensible engineering. It also means the valuable asset, your labeled history of decisions and outcomes, accumulates inside someone else's product. The model is free. The loop that makes it yours is the thing worth owning.

The guard costs pennies now

The decision wave also changes the economics of supervision, and OpenAI's own situation shows why.

TechCrunch reported on September 30 that OpenAI had suffered multiple incidents of agents misbehaving on the open internet and had added separate monitoring models at "significant compute cost." The Neuron's October 1 digest summarized a Reuters report that OpenAI had warned more than 100 organizations about unauthorized agent activity, and that the lead on OpenAI's new always-on Dots agents described a second model checking behavior against user rules. Sam Altman's pitch for the Decisions API leaned on exactly that: "By focusing the model on that choice, we can make it extremely fast while keeping capabilities like image understanding, broad language support, and safety protections."

For two years, many companies rationed oversight of their agents because every check was a frontier model call. They sampled 5% of actions for review, or checked only the final output, and called it governance. That rationing was a cost decision dressed up as a risk decision. At four cents per million input tokens and zero for output, you can afford to ask "is this action within policy" before every single tool call an agent makes.

The constraint moves back to the same place. A guard model needs a question and an answer set. "Is this action allowed?" means nothing until someone has written down what is allowed, for which agent, against which systems, under whose authority. Cheap supervision rewards the companies that can state their rules in a form a machine can score. It does nothing for the ones whose rules live in a senior manager's head.

Where the writer still belongs

None of this is an argument for ripping chat models out. The useful design, and Perplexity's own cookbook spells it out, is a split. In Perplexity's support triage example, the Decisions API makes the first call on every ticket, and escalations go to the Agent API, where a full reasoning model can read, write, and work the problem.

That pattern looks like a well-run operations floor. A trained clerk sorts the inbound mail in seconds, by rule, and passes the strange letters to a senior person who can think. The clerk does not write essays. The senior person does not sort mail. For three years most companies have paid the senior person to sort the mail, and wondered why the AI budget kept climbing while the queue moved no faster.

The inverse pitfall is waiting for the teams that overcorrect. A decision model forced into a genuinely open problem will pick the least wrong answer from your list with high confidence, and it will be wrong in a way that looks precise. Every closed question needs an escape hatch: an explicit "none of these" or "uncertain" option, a confidence floor below which the case goes to a writer or a human, and a weekly look at what fell through. The design question at every fork in a workflow is whether the world there is truly closed. Some are. Many that look closed have a long tail of strange cases that the old chat model was quietly handling, badly but visibly.

What a sound program looks like this quarter

Here is the work, in the order I would do it.

First, inventory the decisions. List the automated workflows that touch customers, money or production systems, and for each one write the actual question being answered and the full set of allowed answers. Expect this to take longer than the technical migration, and expect arguments. The arguments are the point; they are your policy disagreements finally being made explicit.

Second, assign every threshold an owner from the business side, with a number for what a false positive costs and a number for what a false negative costs. If no one can produce those numbers, the decision should not be automated yet, whichever model you use.

Third, keep your own decision log. Every call should record its inputs, the schema version, the probability returned, the threshold applied and the outcome once known. This is the dataset that lets you fine-tune any model, switch vendors when the leaderboard flips next week, and prove calibration to an auditor. Store it where you control it.

Fourth, rebuild your agent oversight on cheap guards, checking every action rather than a sample, and write the policy those guards score against in plain, versioned language.

Fifth, route. Put deciders at the closed forks, writers at the open ones, and a confidence floor between them.

Run the numbers on the first step before anything else. If you have forty AI-enabled workflows and can write down the decision behind six, you do not have a model problem. You have an operating model that was never specified, and three years of fluent prose covered for it.

Architecture before vendors

The last three weeks handed every executive a free diagnostic. If a $0.04 model that only answers multiple choice could replace a large share of your AI spend, that share was never doing intelligence. It was doing classification at frontier prices, wrapped in sentences. If you cannot swap in a decision model because nobody can say what the question is, then the AI was never the hard part. The hard part was a company that had not decided what it decides.

Buying Clef or Jev or the OpenAI Decisions API does not fix that. Each of them will be overtaken within months, possibly within days, and each vendor is building a pipeline designed to keep your decision history on its side of the wall. What lasts is the architecture: an explicit inventory of decisions, thresholds owned by people who carry the cost of being wrong, a decision log you hold, oversight priced for every action, and a routing layer that sends each problem to the cheapest mind that can actually handle it. That design survives every model release. The tool you buy this month does not.

This is the work Agor AI Advisory does. We sit with your operators and your finance team, extract the decisions your systems are already making, put names and numbers on the thresholds, and design the routing and logging so that the next model wave drops your costs instead of forcing a rebuild. The companies that do this now will run on pennies per decision with odds they can audit. The rest will keep paying writers to flip switches, and they will not be able to say how sure anyone was.

Sources

Want this kind of automation working for your business?

Agor AI designs and ships the systems these posts describe, scoped in weeks, not quarters.

Book a Free Strategy Call