← Back to Insights

Insight

The Answer Has No Words

Ariel Agor
•
The Answer Has No Words

Listen · Read by Ariel · click any word to jump

0:00 / —
· loading…

On September 15, 2026, a San Francisco startup called TypeSafe AI released a frontier model that cannot write. You give Jev some application state and a set of questions you have defined in advance. It returns typed answers, each with a calibrated confidence score, and no prose of any kind. Sixteen days later, on October 1, Cloudflare and Amazon's Strands Labs both shipped open-weight models that do the same thing. Three launches in a little over two weeks, from a seed-stage lab, an internet infrastructure company and a hyperscaler's research arm, tell you where the money is. Look honestly at the generative AI business use cases your company has funded since 2023 and most of them turn out to be decisions. Somebody needed a label, a route, a score or a yes. You paid a language model to write a paragraph around that answer, then paid an engineer to dig the answer back out of the paragraph.

I think this is the most useful correction to enterprise AI strategy this year, and it came from a model that answers in numbers.

The sentence was packaging

Think about what happens when a support ticket arrives at a company running a typical 2025 deployment. The ticket goes to a large language model with a prompt that says something like "classify this ticket into one of the following categories and explain your reasoning." The model writes a few hundred tokens. Code parses those tokens, hoping the category name appears in the expected spot and spelled the expected way. If the model wrote "Billing (possibly Refunds)" the parser breaks, a retry fires, and the bill doubles. Somewhere in a dashboard this shows up as a generative AI use case in customer service.

The product of that whole pipeline is one word from a list of twelve. Everything else was packaging, and the packaging was the expensive part. You paid for every output token. You waited for every one of them to stream. Then you built validation, retries and fallback logic to protect your software from the fact that a writing machine sometimes writes something unexpected.

TypeSafe's own documentation is unusually blunt about the alternative. "Jev is the wrong tool for chat, code generation, or anything needing a written explanation," the company says, as quoted in DataCamp's explainer. That is a strange thing for an AI company to put in writing in 2026, when every launch post promises a model that can do everything. It is also the most commercially honest sentence I have read from a lab in months.

Three launches in sixteen days

The details matter here, because the category is new and the claims vary in how much you should trust them.

TypeSafe and Jev

TypeSafe came out of stealth with a $40 million seed round led by DCVC, according to Dealroom's coverage of the launch. The CEO is Diogo Almeida, a former OpenAI researcher and co-inventor of RLHF, the technique behind ChatGPT. His co-founders are Erik Gafni and Sasha Sheng. Almeida's framing of the company fits in one line: "Most intelligence should eventually live inside software, running quietly in the background."

TypeSafe calls Jev a System One model, borrowing Daniel Kahneman's term for fast, intuitive judgment. Per DataCamp, Jev charges $0.042 per million input tokens and nothing for output, answers in 70 to 500 milliseconds, and claims to be 40x to 200x faster and 40x to 400x cheaper than frontier LLMs. The accuracy figure is the one I would put in front of a board. On TypeSafe's own four-workflow benchmark, Jev scored 67.8% against 67.9% for GPT-5.6 Terra. The claim is parity, at a fraction of the cost and latency, which is a far more believable claim than superiority.

The market reaction was immediate. Crypto Briefing reported on September 26 that Vercel saw "more than double its typical paid account sign-ups within 24 hours" of the release, that the launch video passed 40 million views on X in under a week, and that The Information had reported TypeSafe in talks to raise more than $1 billion at a valuation above $10 billion. The seed had valued the company at $200 million. DCVC's James Hardiman told reporters the company is "already generating profits thanks to its low inference costs."

Cloudflare and Clef

On October 1, Cloudflare released the first models it has trained itself. The Cloudflare changelog entry describes Clef (27 billion parameters) and Clef-flash (9 billion parameters) as decision models that "return typed answers with probabilities." Both are on Workers AI, and both are on Hugging Face under Apache 2.0, so you can run them on your own hardware.

Cloudflare's numbers are pointed directly at TypeSafe. Across 43 runs, Clef-flash posted a median latency of 38.8 milliseconds, Clef 209.3 milliseconds, and Jev 524.1 milliseconds. Clef leads 7 of 10 decision benchmarks by Cloudflare's count, including 94.20% on BANKING77, a standard test of sorting banking customer requests into intents. Cloudflare says Clef beats Jev on three of four of Typesafe's own workflow evaluations, covering invoice processing, customer service and security incidents. The models accept up to 64 questions per request, and Cloudflare built them to be drop-in compatible with Jev's System One API. In one example, Clef classified a domain in 2.2 seconds where gpt-oss-120b took 4.7.

The compatibility claim is the important one. Cloudflare looked at a startup's interface two weeks after launch and decided to copy its shape, so a company that builds against that interface can switch suppliers without rewriting the integration. Once a second vendor ships the same API shape, a product has started turning into a category.

AWS and Strands Decider

The same day, AWS's Strands Labs released Strands Decider 2B. Techstrong.ai's October 2 report gives the specifics. The model has about 1.9 billion parameters. Its builders took Qwen3.5-2B-Base, threw away the language-modelling head, and replaced it with a pointer head of roughly one million parameters that picks among the options you supply. It runs on a CPU, a consumer GPU or a Mac. Median latency is 115 milliseconds per question on an Nvidia RTX 3090.

AWS was careful about the limits. Decider 2B cannot write code, summarise documents, hold conversations or handle complex reasoning. On 231 JevBench tasks it scored 72.3% overall, with 100% on easy tasks and 50.5% on hard ones. It is an experimental Strands Labs project, and AWS does not offer it as a production service.

So in sixteen days we got a well-funded startup claiming parity with the frontier, an infrastructure company claiming it is faster, and Amazon releasing a small, honest, local version with its weaknesses written on the label. When competitors this different converge on the same interface that quickly, they are responding to the same pile of demand. The demand is your company's decisions.

Generative AI business use cases, recounted

Pull up whatever list your team made of generative AI business use cases in 2024 or 2025. I would bet heavily on what is on it. Ticket routing. Lead scoring. Invoice coding. Contract clause flagging. Fraud triage. Content moderation. Résumé screening. Churn risk. Product categorisation. Sentiment on reviews. Deciding which agent tool to call next. Deciding whether a draft is good enough to send.

Every item on that list ends in a choice from a known set or a number on a known scale. Hardly any of them end in a document a human reads for pleasure. The "generative" label stuck because the only capable model available was a text generator, and so every problem got phrased as a writing assignment.

Deloitte's State of AI in the Enterprise 2026 report, drawn from 3,235 leaders surveyed in August and September 2025, has a gap in it that I read as this exact mistake. Of the organisations surveyed, 53% said AI had enhanced insights and decision-making, and 66% reported productivity gains. Only 20% said AI had increased revenue, while 74% hope it will. Another 37% said they are using AI at a surface level "with little or no change to existing processes."

My explanation for that gap goes like this. A writing model sits beside the process. It drafts the email a person sends, writes the summary a person reads, suggests the category a person confirms. The person stays in the loop because prose needs a reader. A decision model sits inside the process. Its output is a typed value that software can act on directly, along with a number saying how sure it is. That is the difference between an assistant and a component, and components are what change processes. Companies that bolted a chatbot onto an unchanged workflow got surface-level results because chatbots are a surface.

Where the writing model still belongs

Some real use cases are generative. Drafting a proposal is one. Writing code, explaining a policy change to customers and summarising a 90-page filing for a board member who will never read the original are others. Those stay with the large models, and the large models keep getting better at them. The mistake was sending every task through the same expensive model, which is how a company ends up using a novelist to sort its mail.

The new split looks a lot like Kahneman's. Fast, cheap, bounded judgments go to System One models. Slow, open-ended work that needs reasoning or a written explanation goes to System Two models. An agent built this way uses the small model hundreds of times for every one call to the large model, deciding which tool to use, whether a result passes, whether to escalate, and saves the expensive reasoning for the step that really needs it. AWS listed model routing, tool selection, guardrails and evaluations as Decider's intended uses. Every one of those is a decision your agents are currently paying a frontier model to make.

A probability is a management tool

Most commentary on these launches has focused on speed and price. I think the bigger change is the calibrated confidence score, because it gives executives something language models never gave them, which is a dial.

When a language model writes "this transaction appears to be fraudulent," you cannot do much with the word "appears." You can't set a policy on it or audit it against outcomes, and you can't tune it without rewriting the prompt and hoping. Engineers who asked LLMs to rate their own confidence got numbers that were mostly decoration, because the model was generating a plausible-looking digit with no fixed relationship to how often it was right.

A decision model that returns a calibrated probability is making a promise you can test. If it says 0.9 across a thousand cases, roughly nine hundred of them should be right. You can check that against your own data in an afternoon. Once it holds, the probability becomes something you can manage.

The threshold is the policy

Suppose your accounts payable team runs invoice coding through a decision model. Every invoice comes back with a general ledger code and a confidence. Somebody now has to decide the cutoff. Above 0.97, post automatically. Between 0.80 and 0.97, send to a clerk with the suggested code filled in. Below 0.80, send to a senior accountant. Those three numbers are the policy, and they belong to the controller, not to whoever wrote the prompt.

Every risk appetite statement your company has ever written now has somewhere to live in software. A bank's fraud team can lower the auto-block threshold during a holiday spike and raise it back in January. A hospital's intake desk can be stricter about triage on weekends when fewer clinicians are on shift. The setting can sit in a config file with a change log and an owner, and a committee can argue about it the way it argues about credit limits.

That moves AI governance from the prompt engineer's desk to the operating executive's. I think that shift is overdue and most companies are not ready for it. When your AI component outputs a probability, someone in the business has to own the number that turns it into action. If nobody does, the default is whatever the developer typed during testing.

The escalation queue becomes the product

Thresholds create a second thing worth designing, which is the pile of cases the model was unsure about. In a well-run deployment, that queue holds the hard cases, the ones where your people's judgment actually matters. It is also the best training data you will ever have.

Cloudflare launched Clef alongside a reinforcement learning fine-tuning platform, according to its announcement headline. The loop is simple. Low-confidence cases go to humans, the humans decide, and their decisions train the next version of the model. Over time the band of uncertainty narrows for your particular business, with your categories, your customers and your edge cases. A company that runs this loop for a year owns a model tuned to its own operations that a competitor cannot buy off the shelf. A company that keeps sending everything to a general chatbot owns nothing but a list of past invoices.

Jevons gets the last word

TypeSafe named its model after the Jevons paradox, the nineteenth-century observation that more efficient steam engines raised Britain's total coal consumption, because cheaper energy made new uses of coal worth it. The name is a forecast. When a decision costs a fraction of a cent and returns in a tenth of a second, companies start making decisions they used to skip.

DataCamp's write-up of Jev gives one example: scoring a 50-million-row table of product reviews for sentiment for about $20, against thousands of dollars with token-billed LLMs. At thousands of dollars, that job needed a business case and a sign-off. At $20 an analyst runs it on a whim, and so most of the latent demand sits in decisions your company never made at all.

Say a retailer with 400,000 SKUs has always reviewed product descriptions for compliance problems by sampling 2% of them each quarter, because human review is slow and LLM review was too expensive at full scale. With a decision model it can check every SKU every night. The same goes for every contract in the repository, every support transcript and every vendor invoice against the purchase order. Work that used to be audited by sample can now be checked in full, and a full check catches different problems than a sample does.

Pricing changes too. Jev bills nothing for output, Clef returns a single typed value, and Decider runs free on hardware you own. Once output is free, the meter runs on what you send in and on how many questions you ask, and the economics favour asking many narrow questions over a few long open ones. Engineering teams used to writing long, rambling prompts will have to learn to write schemas, which is a different and more disciplined skill.

Where this goes wrong

I would not let anyone rebuild a process on these launches without reading the fine print, and the fine print is substantial.

Start with the benchmarks. Most of the published numbers come from the vendors, and some come from the vendors' competitors. Cloudflare benchmarked against Jev. TypeSafe benchmarked against GPT-5.6. JevBench, which Techstrong used to report Decider's results, carries the name of the product it was built around. None of that makes the numbers false. It means your own data is the only benchmark you should trust, and you should run it before you sign anything.

Hard cases are where these models struggle. Strands Decider got every easy JevBench task right and about half of the hard ones. That profile suits triage well, because the easy cases are most of the volume and the hard ones should go to a person anyway. It would be dangerous for anyone who reads "72.3%" as a uniform accuracy and sets a single threshold across the whole distribution. Calibration is what makes the hard half safe, and you have to check calibration on your own cases.

TypeSafe's valuation adds another risk. A company that went from $200 million to talks at more than $10 billion in under two weeks will be under heavy pressure to grow into that number, and pressure like that changes pricing, terms and roadmaps. Cloudflare's decision to copy Jev's API and release Apache 2.0 weights is the best protection a buyer has against it. Build against the interface, keep an open-weight model qualified as a fallback, and you keep the option to walk away.

The last risk sits inside your own organisation. Decision models are easy to drop into a workflow without telling anyone, because they make no visible noise. A chatbot that writes something offensive gets screenshotted. A classifier that quietly sends 4% of a demographic's loan applications to a slower queue does not. The thresholds and queues that make these models manageable also make their failures quiet. You need someone watching the distributions, and that person needs authority to stop the line.

Rebuild the list before you buy anything

The practical work starts with a different inventory. Most companies hold a use case list organised by department and tool, such as "marketing uses Jasper" or "support uses a Zendesk bot." Throw it out and list your decisions instead.

For every recurring decision in the business, write down the inputs it uses, the allowed answers, how many times a day it happens, how long it waits for a human today, and what a wrong answer costs. Most companies that do this find hundreds of them. Then sort them. Decisions with a fixed answer set, high volume and a cheap cost of error go to System One models first. Decisions with a fixed answer set and an expensive cost of error go to System One models with conservative thresholds and a human queue. Decisions that need explanation, negotiation or new text stay with System Two. Decisions that are rare and expensive stay with people, and you should be suspicious of any vendor who tells you otherwise.

Next, give each threshold an owner. The owner should be a person with operating responsibility for the outcome, not someone from IT. Their name goes on the config file and their review cadence goes on the calendar.

Finally, treat the escalation queue as an asset. Measure how much smaller it gets each month. If it is not shrinking, you are not learning, and the deployment is a cost centre with good latency.

Companies that skip this work will still buy decision models, because their developers will adopt them for speed and price alone, the way they adopted every cheap API before. They will get faster software. They will not get the decision layer that compounds over time, because nobody wrote down which decisions mattered, nobody owns the thresholds and nobody trains on the queue. Two companies can pay the same vendor the same amount and end up eighteen months apart.

Architect the decisions, then pick the model

The September and October launches handed every executive a test. You can read them as a cheaper API, swap a few endpoints and log the savings. Or you can treat them as the first chance to rebuild your operations around decisions that have owners, thresholds and a learning loop, in models whose weights you can hold.

The first reading saves money. The second gives you a business that makes ten times as many good decisions per day and knows how confident it was in each one. Getting there takes architecture work, which means mapping decisions across functions, setting the governance for who owns which probability, designing the human queues and qualifying open-weight fallbacks before a vendor's valuation turns into your price increase. A procurement team cannot do that work and a software vendor will not do it for you, since a vendor's incentive is to sell you its model for every decision, including the ones it handles badly.

Agor AI Advisory does this work with leadership teams. We build the decision inventory with you, split it between System One models, System Two models and people, put named owners on the thresholds, and design the loop that turns your hardest cases into a model your competitors cannot buy. The models launched in the last three weeks. The companies that map their decisions this quarter will set the cutoffs that everyone else ends up copying.

Sources

Three decision models, three different kinds of claim

The post argues that the new decision models converged on one interface but make claims that are not comparable with each other. After 15 seconds the reader should see that each vendor measured something different, so only the reader's own data can settle the choice.

  • Every number in this table comes from a vendor or a vendor's competitor. Your own data is the only benchmark that counts.
  • Decider's 100% on easy tasks and 50.5% on hard ones are the same model. A single accuracy figure hides the split.
  • Three vendors converged on one interface in sixteen days, so the integration is portable and the choice of model can still change.
Reported latencyAccuracy claimAvailability and price
Jev (TypeSafe)Parity with a frontier LLM at a fraction of the cost, but the accuracy benchmark is the vendor's own and its latency is disputed by a competitor.524.1 ms median in Cloudflare's 43 runs; TypeSafe claims 70 to 500 ms67.8% on TypeSafe's four-workflow benchmark vs 67.9% for GPT-5.6 TerraSystem One API; $0.042 per million input tokens, nothing for output
Clef and Clef-flash (Cloudflare)Fastest and open-weight, so it is the easiest to switch to, but the benchmarks were chosen and run by Cloudflare against a rival.209.3 ms (Clef, 27B) and 38.8 ms (Clef-flash, 9B) median across 43 runsLeads 7 of 10 decision benchmarks by Cloudflare's count, including 94.20% on BANKING77Workers AI and Hugging Face under Apache 2.0; drop-in compatible with Jev's API
Strands Decider 2B (AWS Strands Labs)Local and honest about its limits, but it gets about half of the hard cases wrong and carries no production support.115 ms median per question on an Nvidia RTX 309072.3% on 231 JevBench tasks: 100% on easy, 50.5% on hardRuns on a CPU, consumer GPU or Mac; experimental, not a production AWS service

Source: Figures as cited in the post from the Cloudflare changelog (Oct 1, 2026), Techstrong.ai (Oct 2, 2026) and DataCamp's Jev explainer (Sept 2026). All numbers are vendor-reported. · verified · as of 2026-10-06

Want this kind of automation working for your business?

Agor AI designs and ships the systems these posts describe, scoped in weeks, not quarters.

Book a Free Strategy Call