JOURNAL
Three Decision Models Shipped in One Week — None of Them Wins Outright
Cloudflare's Clef/Clef-flash and Amazon's Strands Decider both launched October 1st aiming at TypeSafe's Jev. I lined up the real benchmark, latency, and pricing numbers — the category is real, but picking a winner depends entirely on which decision you're automating.

On this page
Two weeks ago I wrote about Jev, TypeSafe’s bet that a huge chunk of agent decisions don’t need a chat model at all — just a calibrated yes/no, a score, or a pick from a fixed list. I called it a System 1 for agents. I did not expect the rest of the industry to show up this fast.
October 1st, two separate teams shipped direct answers to Jev on the same day. Cloudflare released Clef (27B, built on Qwen3.8) and Clef-flash (9B, built on Qwen3.5), both open-weight under Apache 2.0, both serving on Workers AI. Amazon’s Strands Labs released Strands Decider 2B — also Apache 2.0, also built on Qwen3.5, but a completely different trick under the hood. Three “decision model” releases, two architectures, one week. That’s not a coincidence, that’s a category forming in real time.
So I did the thing I actually wanted to do with the Jev piece: put real numbers from all three side by side and see who actually wins.
They’re not solving the same problem the same way
Jev is a proprietary hosted model — no published weights, 64K context, 32K of that reserved for state plus your longest question. You call an API, you get a typed answer back.
Clef and Clef-flash are prefill-only, non-autoregressive. The model never generates tokens one at a time — it reads your context once and returns a probability distribution over typed outputs in a single forward pass. That’s architecturally why they’re fast: no autoregressive loop, no sampling, no stop tokens to wait for.
Strands Decider does something stranger and, honestly, more interesting. AWS took Qwen3.5-2B-Base, cut off the language-model head — the part that would normally generate the next token as text — and bolted on a pointer head with roughly 1 million parameters, plus a rank-16 LoRA adapter on the backbone. The pointer head doesn’t generate anything. It scores the answer options you hand it directly. You’re not asking the model to write “yes” — you’re asking it to point at option 0 out of the two you gave it, with a calibrated confidence attached.
Three different ways to avoid the thing that makes chat models slow: writing text you’re going to throw away anyway.
The numbers, lined up
| Jev | Clef | Clef-flash | Strands Decider 2B | |
|---|---|---|---|---|
| Base | proprietary | Qwen3.8-27B | Qwen3.5-9B | Qwen3.5-2B |
| License | closed | Apache 2.0 | Apache 2.0 | Apache 2.0 |
| Median latency | 524.1 ms | 209.3 ms | 38.8 ms | ~115 ms (RTX 3090) |
| Price (input, /1M tok) | $0.042 | $0.240 | $0.090 | free, self-hosted |
| BANKING77 (macro-F1) | 79.74 | 94.20 | 90.93 | — |
| GPQA Diamond | 78.3 | 48.0 | — | — |
| MMLU-Pro | 82.7 | 65.9 | — | — |
| JevBench hard tier | — | — | — | 0.505 |
A few things jump out, and none of them are “buy Clef” or “buy Jev.”
Jev is still clearly ahead on anything that smells like general reasoning — GPQA Diamond, MMLU-Pro, When2Call accuracy (80.97 vs Strands’s “hard tier” score of 0.505, which AWS’s own writeup is honest enough to call “basically a coin flip”). If your decision is “should the agent escalate this to a human” or “is this action safe to take,” you want the model that’s actually reasoning about the situation, not pattern-matching against a fixed label set. That’s still Jev’s lane.
Clef absolutely dominates intent classification and routing — BANKING77, CLINC150+OOS, tool-call selection. These are closed-set, pattern-heavy decisions where you’re not asking the model to reason, you’re asking it to recognize. Clef-flash is 13x faster than Jev at 2.1x the price; full Clef is 2.5x faster at 5.7x the price. If your bottleneck is “which of these 40 intents is this ticket,” Clef-flash is the obvious pick and it’s not close.
Strands Decider’s real pitch isn’t the benchmark numbers — it’s that AWS published everything. Weights, training data, training scripts, the works. Clef and Clef-flash are open-weight in the sense that you can download and run them, but Cloudflare hasn’t published the training pipeline or data. Strands is the one you can actually audit and reproduce from scratch, if that matters to your compliance team more than squeezing out another 5 points of accuracy.
The part the launch posts don’t lead with
Strands Decider’s reference server binds to localhost with no authentication by default. That’s a reasonable default for a local dev loop and a genuinely bad one if someone copies the quickstart into a shared environment without reading past step 3. If you’re self-hosting this — and self-hosting is the entire reason to pick it over Jev — put it behind your own auth layer before anything touches it. This is the same category of mistake as shipping a database with no default password: fine until it isn’t.
The other thing nobody’s benchmark table captures: these are all narrow models with hard failure modes outside their training distribution. Clef-flash’s CLINC150+OOS score (66.77, well behind Clef’s 97.43) is the tell — it’s good at picking from known categories and noticeably worse at recognizing “none of the above.” If your agent’s decision space can genuinely surprise it, a fast narrow model that’s confidently wrong is worse than a slow general one that hedges.
Where I’d actually put these
I’d run Clef-flash in front of a support or ticket-routing pipeline where the intent set is fixed and well-understood — the latency and cost math are too good to ignore for a closed-set problem. I’d keep something Jev-shaped (or a small general model) behind any decision that gates an irreversible action — refunds, deletions, anything you can’t undo. And I’d only reach for Strands Decider if “we can retrain this ourselves, air-gapped, on hardware we own” is an actual requirement, not a nice-to-have — because the accuracy ceiling right now is clearly below both Jev and Clef on anything that isn’t pure classification.
Three launches, one week, and the honest conclusion is the boring one: this isn’t a model you pick, it’s a slot in your architecture you fill differently depending on what’s actually at stake if the model is wrong.



Discussion
Comments are reviewed before publication. Your email is kept private.