
Chat models have been better than most people at following instructions for years. So why is so much business software still held together by hand-written rules?
That question is what pushed Diogo Almeida to start TypeSafe AI. He worked at OpenAI on the research behind ChatGPT. His answer is a model called Jev, and it does something no chatbot does: it never writes a sentence.
What Jev actually does
You send Jev some state, either plain text or structured data, along with a set of typed questions. It answers all of them in one pass.
Each answer is a value from a list you defined in advance, such as a choice from fixed options or a score. Each one comes with a probability and a confidence value.
TypeSafe describes it as "a frontier-intelligence function call": messy input in, typed decisions out.
Here is an example. Say a support ticket arrives. You ask Jev whether it is urgent and which team should handle it. It returns an answer to each, with a probability attached.
Your code then does the rest. If the confidence is high, it acts. If not, it sends the ticket to a person.
Why it is fast and cheap
A normal language model writes one token at a time, and each token depends on the one before it. Jev produces all of its outputs in a single parallel pass.
Because the answers come from a schema you set, there is nothing to parse or validate afterward. TypeSafe says that is why the model can never make a type error.
TypeSafe's own comparison looks like this:
Typical LLM | Jev | |
|---|---|---|
Output | Generated text | Typed values with probabilities |
Sampling | One token at a time | All outputs in parallel |
Input price | $0.20 to $10 per million tokens | $0.042 per million tokens |
Output price | About 5x the input price | Free |
Response time | 3 to 329 seconds (frontier models) | 70 to 500 ms |
The context window is 32,000 tokens, so this is not a tool for stuffing in whole codebases.
The model is trained with a method TypeSafe calls Reinforcement Learning for Calibrated Decisions. The goal is calibration: higher confidence should mean higher accuracy.
The names have a story too. "System One" comes from Daniel Kahneman's Thinking, Fast and Slow. "Jev" is a nod to William Stanley Jevons, on the bet that cheaper intelligence leads to more demand for it.
Where it fits
TypeSafe calls the use case "smart if-statements." Classify, route, score, extract or branch, in places where a hand-written rule is too brittle.
It can also judge the output of other models, which is useful for guardrails. LangChain has shipped an integration that uses Jev for model routing and for screening tool calls before an agent runs them.
TypeSafe's demos show the speed. In one, Jev plays Doom from a text description of the game state, at about 10 queries a second. TypeSafe puts that at roughly $7 an hour.
The headline numbers, and who ran them
The homepage claims Jev is 193.6 times faster and 444.6 times cheaper. Those figures come from TypeSafe's own workflow tests.
The company is open about the limits of that test. The four workflows were built by its own team. There is no answer key. Instead, the reference answers are the average of what GPT-6 Astra and Fable 5.1 predicted.
TypeSafe says these results are likely at the high end of real-world gains. Elsewhere it quotes a range of 40 to 200 times faster. It also admits it cannot prove its pricing is not subsidized.
What outside users found
The early reports back up the direction, with smaller numbers.
Vercel added Jev to its AI Gateway. It says nearly 13% of its paid teams used it within 24 hours, about twice the share of the GPT-5.6 family.
A Vercel engineer found a safety classifier ran 5 to 18 times faster than the LLM it replaced.
The CTO of Bryo AI found Gemini slightly more accurate on email classification, but 10 to 20 times more expensive. He liked that Jev returns a real probability.
An analysis of 12,759 launch tweets put the median reported speedup at 7x, the median cost saving at 30x, and the median latency at 76 ms.
So the gains look real. They also look a lot smaller than the headline.
The limits
Jev cannot write, so it cannot replace a model that writes code or drafts text. It sits next to one.
"Zero hallucinations" needs reading carefully. Jev cannot return an invalid type. It can still pick the wrong valid answer, and a developer on Hacker News made exactly that point.
Armin Ronacher, CTO of Earendil, told TechCrunch that the design hands part of the problem to the user. You have to decide whether a 50% probability is worth acting on.
TypeSafe's own documentation for jev-1.13 lists weak spots: counting, arithmetic and date comparison, plus lower accuracy on large, noisy input. Its advice is to keep the maths in code. It also suggests pinning a specific version instead of the moving latest alias.
The company and the money
TypeSafe is based in San Francisco and was founded in 2024 by Almeida, Erik Gafni and Sasha Sheng. It came out of stealth on September 15, 2026, with a $40 million seed round led by DCVC.
Forbes reported a valuation of about $200 million. TypeSafe has not confirmed that figure.
The Information later reported that the company was in early talks to raise more than $1 billion at a valuation above $10 billion. That was a report about talks, not a finished deal, and TypeSafe has not confirmed it.
Should you try it?
If your product makes LLM calls that are really classification or routing, Jev is worth testing on your own data.
TypeSafe offers an adapter that runs other models against the same schema, which makes a side-by-side comparison easy. Check whether the confidence scores match the accuracy you actually see.
Do not expect it to replace your chat model. The idea is that the chat model talks to people, and Jev makes the small decisions in between.
Sources
Introducing System One Models & Jev, TypeSafe AI, September 15, 2026
TypeSafe AI Releases Jev, InfoQ, October 1, 2026
TypeSafe AI emerges from stealth with $40M, Business Wire, September 15, 2026
Comments (0)
Join the discussion by logging into your account.
No comments yet. Be the first to comment!