A customer writes in: “I’ve been trying to connect my payment account for three days and it keeps failing. I’m losing orders. Please fix this now.” Feed that into a normal LLM and you get back something like: “Based on the message provided, this appears to be a Billing issue because the customer is experiencing a payment-related error…” Your routing code doesn’t want that paragraph. It wants one value: billing. Everything before the colon was generated for no one.
TypeSafe AI shipped a model on September 15, 2026 that refuses to write that paragraph at all. It’s called Jev, and it’s the first release in what they’re calling a “System One Model” line — a direct nod to Kahneman, and Simon Willison picked up on that framing immediately in his September 21 write-up, calling it “a new shape of LLM.” Jev doesn’t chat, doesn’t summarize, doesn’t write code. You hand it unstructured state and a pre-declared set of valid answers, and it hands back a typed decision with a number attached to how sure it is.
The company behind it, briefly
TypeSafe AI is a two-year-old SF startup, not a rebrand of the old Scala/Akka “Typesafe” that became Lightbend — different company, unfortunate name collision. Founders are Diogo Almeida, who spent about four years at OpenAI and is credited as a co-inventor of RLHF, along with Erik Gafni and Sasha Sheng. $40M seed, led by DCVC, at roughly a $200M valuation. Coverage landed fast and from places that don’t usually agree on what’s worth writing about — TechCrunch, The Register, Tom’s Hardware, Forbes, and a Wikipedia page already exists for the model itself. That spread is usually a decent signal something real happened, as opposed to a blog post nobody outside one Discord server read.
Three ways to ask a question, and nothing else
Jev ships exactly three “primitives,” and the constraint is the feature. Choice picks from a fixed list — up to 255 options — and returns a probability for every option plus a separate confidence score derived from the shape of that distribution. Score rates something on an ordinal scale, two to ten levels, say severity from low to critical, and returns the score, the full distribution, and confidence. Noul is the odd one out by name — it’s TypeSafe’s own coinage, short for Bernoulli, and it’s just yes/no: a single probability between 0 and 1, no separate confidence field because with one number there’s nothing left to derive.
That’s it. No free-form field, no “explain your reasoning” option. If your use case doesn’t fit one of three shapes, Jev is the wrong tool, and that’s stated as a design choice rather than apologized for.
Why the distribution matters more than the answer
Take that billing ticket again, run as a three-way Choice between billing, engineering, and sales. One outcome: billing: 98%, engineering: 1%, sales: 1%. Auto-route it, no human needed. A different outcome, same top answer: billing: 52%, engineering: 46%, sales: 2%. Billing still wins, but a system that only reads the winning label can’t tell these two cases apart — and they’re not the same case. The second one is a coin flip wearing a label. If your routing logic only ever sees the string "billing", you’ve thrown away the one piece of information that would have told you to hold the ticket for a person instead of auto-closing it.
This is the part worth sitting with if you’re on the business side rather than staring at API docs all day: the value here isn’t that Jev is right more often. It’s that when it’s not sure, it says so in a format code can act on, instead of sounding equally confident every time the way most chat-tuned models do.
RLCD, and a naming collision nobody’s addressed
TypeSafe trains Jev with a method they call RLCD — Reinforcement Learning for Calibrated Decisions — optimizing directly for calibration rather than for a human rater’s stylistic preference. Calibration here means something specific: if the model says “90% confident” on a hundred separate calls, roughly ninety of those should turn out correct. A model that’s right 95% of the time but claims 100% confidence on everything is worse for automation than one that’s right 93% of the time and actually knows which 7% it’s shaky on, because only the second model tells you when to escalate.
Here’s my one complaint, and I don’t think it’s minor: RLCD already means something in this field. “Reinforcement Learning from Contrastive Distillation,” Yang, Klein, Celikyilmaz, Peng, and Tian, published to arXiv in 2023 and presented at ICLR 2024, is a real, citable, prior technique with the exact same three-letter name. I found nothing on TypeSafe’s blog or docs acknowledging the earlier paper. Maybe it’s coincidence, maybe it’s an internal team that didn’t run a literature search before naming their loss function. Either way, if you’re reading a paper or a Slack thread that says “RLCD” from here forward, check which one is meant before you nod along.
Where this actually sits in an agent stack
The pitch that makes Jev more than a curiosity is where TypeSafe wants it deployed: not instead of your LLM, but wrapped around it, absorbing the small decisions that don’t deserve a full reasoning pass.
Model routing — classify task difficulty first, send simple extraction to a cheap model, keep the expensive reasoning model for what actually needs it. Tool risk gating — read a coding agent’s proposed tool call and classify it read-only, reversible, or destructive before it executes. LangChain built exactly this as AutoModeMiddleware in a new langchain-typesafe package, wired up in GitHub PR #40556 by maintainer Sydney Runkle — worth flagging that as of this writing the PR is still open, not merged, so treat it as alpha rather than something to put in production this week. Agent supervision — check whether an agent’s output actually satisfies the original ask, or whether it’s spinning in a loop. RAG filtering — after vector search returns candidate chunks, score which ones actually support an answer before they burn tokens in the main model’s context window.
None of that needs a paragraph back. It needs a label and a number, fast, cheap, and repeated thousands of times a day — which is a different job than the one people usually hire an LLM to do.
Where I’d actually push back
I’ve spent enough time lately writing about phase-scheduled multi-agent systems and per-subagent effort dials to recognize the shape of this pitch: it’s another way of admitting that routing everything through one expensive general-purpose model was always a waste, and that most of what an agent loop needs isn’t intelligence, it’s a fast, honest gate. Jev is genuinely new in the “typed probabilistic output as the product” framing. It is not new in the underlying idea — small classifiers gating big models have existed since long before anyone called an LLM an agent. What’s actually new is that TypeSafe is selling the gate itself as the product, with a calibration guarantee attached, instead of leaving every team to hand-roll their own confidence heuristics on top of a generic model’s raw logits.
Whether the calibration claim holds up under someone else’s eval suite, not TypeSafe’s own benchmark page, is the thing I’d want to see before wiring this into anything that touches money or an angry customer.