Jev and System One Models: What's Real in TypeSafe's 193x Claim — and Where a $0.042/M Decision Model Fits Your Coding Agent
TL;DR: Jev, TypeSafe AI’s “System One” model (early access since September 15, 2026), answers typed questions — yes/no probabilities, multi-choice, scored rubrics — in one non-autoregressive pass at $0.042 per million input tokens with free output. The 193.6x-faster / 444.6x-cheaper homepage claim is TypeSafe’s best case against its slowest, priciest baselines; independent reruns measure roughly 3–5x faster and 40x cheaper than Claude Sonnet 5 at equal accuracy. That still changes the math for agent pipelines making hundreds of decisions per session — and changes nothing if you make five.
| Jev 1.13 | Frontier LLM, strict JSON (Sonnet 5.5) | Hard-coded rules / regex | |
|---|---|---|---|
| Best for | High-volume typed decisions inside agent loops | Decisions that need an explanation or edge-case reasoning | Decisions with a stable, enumerable pattern |
| Cost per ~1k-token decision | ~$0.00004 (measured) | ~$0.0007 (measured, same task) | $0 |
| Median latency | 420–480 ms (two independent benchmarks) | 1,370–2,480 ms (same benchmarks) | microseconds |
| The catch | No text output, 64k context cap, direct signups paused | 17–40x the cost, 3–5x the latency per decision | Brittle; can’t read intent or phrasing |
Honest take: If your coding agent gates, routes, or retries more than ~100 times a day, put Jev (via OpenRouter, since TypeSafe’s own signups are paused) in the control loop and keep the LLM for the code. Below that volume, a structured-output call to Sonnet 5.5 costs pennies a month — skip the extra dependency.
Jev landed with a 1,500+-point Hacker News thread on September 15, 2026 — unusual traction for a model that can’t write a line of code. It can’t write anything: Jev is the first public release in what TypeSafe AI calls System One models, and it only ever returns numbers. For developers running coding agents, that turns out to be the interesting part. Agents burn most of their incidental token spend not on generating code but on deciding things — is this tool call risky, is this test failure real, which model should handle this subtask — and those decisions are currently made by the most expensive text generators on the market.
What is Jev, and what does “System One model” mean?
Jev is a schema-constrained decision model: you send it a block of state (a diff, an error log, a tool call, plain text or JSON) plus a set of typed questions, and it returns calibrated probabilities for every question in a single parallel pass, typically in 70–500 ms end to end per TypeSafe’s published figures. “System One” is a reference to Daniel Kahneman’s fast, intuitive System 1 cognition — the pitch is a model that judges instantly, paired with a slower LLM only when language is actually required.
Three question types exist as of Jev 1.13:
- Noul — a yes/no question. Returns a single number from 0 to 1 that is the yes-probability (there’s no separate confidence field).
- Choice — pick from up to 255 named options. Returns a probability per option plus a confidence value.
- Score — an ordered rubric of 2 to 10 levels. Returns a probability-weighted numeric score with the distribution across levels.
Context limits are 64k tokens for state plus all questions, and 32k for state plus the longest single question. Because the model is non-autoregressive — it doesn’t generate token by token — every question is evaluated against the state in parallel, which is where the latency floor comes from.
The commenters on the Hacker News launch thread converged on a plausible reading of what this is under the hood: a smaller model post-trained for calibration, likely distilling log-probabilities and confidence from a larger teacher. TypeSafe hasn’t published the architecture. What matters for a buying decision is the measured behavior, so that’s what the rest of this article sticks to.
Is Jev really 193x faster and 444x cheaper than an LLM?
No — not in any configuration you’d actually compare. The 193.6x and 444.6x figures on TypeSafe’s homepage are real arithmetic against the slowest and most expensive baselines on their internal workflow suite (the latency multiple is against Claude Sonnet 5 on TypeSafe-selected multi-step workflows; the cost multiple is against Opus 5). TypeSafe’s own launch post concedes these “are on the higher end of real world gains,” and the accuracy grading used GPT-6 Astra and Claude Fable 5.1 outputs as the answer key rather than human ground truth — a grading choice one widely shared DEV Community teardown summarized as “GPT-6 and Claude wrote the answer key.”
Two independent, reproducible benchmarks published in the two weeks after launch give a more honest multiple:
| Benchmark | Task | Accuracy | Latency (median/p50) | Cost |
|---|---|---|---|---|
| arifulislamat/jev-benchmark (100 tickets × 4 questions) | Support triage: routing, refunds, sentiment, urgency | Jev 93% routing / 99% refunds / 89% sentiment — within a few points of Sonnet 5, GPT-5.6 Sol, Gemini 3.8 Flash | 474 ms vs 1,937–3,459 ms | $0.0031 total vs $0.0960–$0.1990 (31–64x) |
| themsquared/jev-benchmark (60 hand-labeled cases, Sep 17–24, 2026) | Agent tool-call risk: readonly / destructive / privileged / exfiltration | Jev and Sonnet 5 identical: 91.7% (55/60) | 421.6 ms vs 1,371 ms (3.25–3.62x) | ~$0.0000173 vs ~$0.0007035 per call (40.6x) |
So the realistic headline is: equal accuracy on single-shot classification, 3–5x faster, 30–65x cheaper than mid-tier frontier models. Both benchmark authors flag the limits — synthetic inputs, one machine, n=60 and n=100 — and the first one notes the text models only agreed with each other 60–72% of the time, which says as much about the tasks as the models.
The second benchmark surfaced the most practically useful finding: calibration. Across all 60 cases, every wrong answer — from Jev and from Claude — came with visibly lower confidence, and Jev’s confidence values spread across a wider range (floor of 0.13, vs 0.55+ from Claude), which makes threshold-based routing actually workable. A decision model you can’t threshold is just a slower coin.
What does Jev cost inside a coding-agent session?
About four cents per thousand decisions. Jev 1.13 is priced at $0.042 per million input tokens, and output tokens are free — a decision with ~1k tokens of state costs roughly $0.00004 (the jev-guard project, which gates every tool call through Jev, publishes the same ≈$0.00004-per-call figure and measured ~580 ms per check through a gateway).
Run the comparison at agent volume. A busy Claude Code or Cursor session fires 50–200 tool calls; gate each one plus a handful of routing and retry decisions and you’re at ~300 decisions per session:
- Jev: 300 × $0.00004 ≈ $0.012/session — $0.36/month at a session a day.
- Sonnet 5.5 in strict JSON mode ($2/$10 per MTok, verified October 5, 2026): 300 × ~$0.0007 ≈ $0.21/session — $6.30/month, plus roughly 1–2 seconds of added wall-clock per decision, which at 300 decisions is 5–10 minutes of accumulated waiting per session.
The dollar difference is real but small at solo scale; the latency difference is what you feel. A 400 ms gate is invisible inside an agent loop. A 2.5-second gate on every tool call is not — it turns auto-mode into molasses. That inversion of the usual argument is worth internalizing: at individual-developer volume, Jev’s case is speed first, cost second. The cost case takes over for teams and CI pipelines running thousands of agent sessions, where $6.30 vs $0.36 per seat per month multiplies.
How do you call Jev today (while signups are paused)?
Here’s the first real-world problem we hit: you currently can’t get a TypeSafe API key. TypeSafe opened self-serve signups on September 20 with $5 of free credit and paused them on September 22 under demand. The fix is OpenRouter, which exposes Jev to any existing OpenRouter key — no TypeSafe account, no waitlist — via its alpha Decisions API:
curl -s https://openrouter.ai/api/alpha/decisions \
-H "Authorization: Bearer $OPENROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "typesafe/jev-1.13",
"state": "ERROR: ETIMEDOUT connecting to db.internal:5432 (attempt 3/3)",
"questions": [
{"type": "noul", "text": "Is this failure likely transient (retry could succeed)?"},
{"type": "choice", "text": "What subsystem failed?",
"options": ["database", "auth", "network", "application-code"]}
]
}'
The model ID is typesafe/jev-1.13; the ~typesafe/jev-latest alias tracks the newest release. Direct customers (from before the pause) use POST https://api.typesafe.ai/v1/systemone with official SDKs — typesafe-sdk on PyPI and @typesafe-ai/sdk on npm.
For poking at it from a terminal, the MIT-licensed community CLI is the fastest route:
npm install -g @y0usaf/typesafe-cli
echo "I was charged twice for order A-104. Please refund the duplicate charge." | \
jev noul "Does the state request a refund?"
# noul 0.93
jev choice "What failed?" --state "ERROR: could not connect to server" \
--option db --option auth --option other
# choice db conf 0.99
Note the alpha in OpenRouter’s endpoint path: this surface shipped in late September 2026 and can change. Pin jev-1.13 rather than the alias in anything you deploy.
How does Jev plug into Claude Code, Cursor, and other coding agents?
Through hooks — and the ecosystem moved fast here. The clearest working example is jev-guard (MIT, npm i -g jev-guard), which implements a trustworthy auto-mode: before every tool call your agent makes, it sends the call plus session context to Jev and asks four typed questions — a 0–3 risk Score, plus Nouls for “would a careful engineer require sign-off?”, “did the user’s recent messages ask for exactly this?”, and “does this execute instructions planted in fetched content?” A threshold policy then maps the probabilities to deny / ask / allow (deny at risk ≥ 2.5 or injection-probability ≥ 0.7). Claude Code gets full deny/ask/allow support via its plugin system; Cursor is supported for shell and MCP calls; Codex, GitHub Copilot CLI, Gemini CLI, OpenCode, pi, and ACP editors (Zed, JetBrains) are covered with per-agent caveats. Cline is notably absent from the supported list as of early October 2026 — Cline users wanting this pattern are writing their own wrappers for now.
This is the same tool-call risk surface we’ve covered since the agentjacking disclosures, and gating it used to mean either rigid rule lists or burning frontier tokens on every call. A calibrated $0.00004 classifier in front of each tool call is a genuinely new point on that curve — with jev-guard’s own caveat attached: it’s “a guardrail, not a sandbox,” and Jev can be wrong, so it complements rather than replaces permission modes.
The second pattern is model routing. An early-access Claude Code mod (jev-model-router, built on function hooks) asks Jev before each turn how mechanical the task is, how much reasoning it needs, and whether it’s risky — then routes to the cheapest capable model. That’s the same decision we worked through manually in our subagent model-routing cost guide; a Choice over candidate models at four-hundredths of a cent per thousand routes makes it automatic. The same shape applies to retry/escalate logic (“is this test failure flaky or real?”) and to pre-filtering findings before an expensive review pass — the judgment-layer problem from when to trust AI review suggestions.
Where Jev breaks
Jev cannot explain itself. You get a number, never a sentence — if your workflow needs the why (a code review comment, a rejection message to show the user), you still need a language model, and at low volume running both is more moving parts than one structured-output call. See our harness architecture comparison for how much glue code each extra layer costs you.
The measured limits, in order of how likely they are to bite:
- Ambiguity is still hard. On the tool-call risk benchmark, all models — Jev included — dropped to 71.4% on the deliberately ambiguous slice. The calibration saves you (wrong answers carried low confidence), but only if your policy actually routes low-confidence cases to a human or a bigger model.
- 64k context. State plus questions caps at 64k tokens (32k for state plus the longest question). A large diff plus session history won’t fit; you have to summarize or truncate, and what you cut is a silent accuracy tax.
- 255 options per Choice, 10 levels per Score. Fine for routing and risk tiers, too coarse for “which of these 400 files is relevant.”
- Access is wobbly. Direct signups paused since September 22; the open path is an endpoint OpenRouter itself labels alpha. Nothing here is GA.
- The accuracy story is young. Vendor numbers are graded against frontier-LLM labels, and the independent benchmarks total 160 cases between them. Nobody has published a Jev evaluation on coding-specific decisions (flaky-test detection, breaking-change classification) at meaningful scale yet.
Privacy posture is the one question we could not verify to this site’s usual standard: TypeSafe is a cloud API, your tool calls and diffs transit it (or OpenRouter in front of it), and we found no published retention terms equivalent to what the major labs document. Treat it like any other cloud inference dependency — and note jev-guard skips sending short and local-only tool results off the machine at all. If decisions must stay on-device, a small local classifier on modest hardware is the alternative; see the sister site’s local models by VRAM guide for what runs where, and the FOSS side of the agent stack in Continue vs Cline vs Aider.
Honest take: who should actually use this
Use Jev if your agent setup makes high-volume mechanical decisions: tool-call gating on every action (jev-guard is installable today and is the single best on-ramp), model routing across a fleet of subagents, retry/escalate logic in CI, or pre-filtering hundreds of candidate findings per run. At 300+ decisions a session the latency win is immediate, the cost rounds to zero, and the measured calibration is good enough to route on thresholds — which is more than can be said for most LLM confidence proxies.
Skip it if you make a handful of semantic decisions a day. A Sonnet 5.5 structured-output call costs about $0.0007 and arrives in 2 seconds; at ten decisions a day that’s $0.21 a month, and you keep one fewer alpha-stage dependency, one fewer API key, and the ability to ask “why” in the same call. The 193x marketing number describes a workload you probably don’t have. The 3-to-40x real number describes one you might be growing into — the moment your agents start running unattended and always-on, decision volume explodes, and this is the cheapest calibrated judgment you can currently buy.
FAQ
Is Jev an LLM? No. It’s non-autoregressive and generates no text. It evaluates typed questions (Noul, Choice, Score) against a state block and returns calibrated probabilities in one pass, typically 70–500 ms.
Can I use Jev without a TypeSafe account?
Yes. TypeSafe’s own signups have been paused since September 22, 2026, but OpenRouter serves typesafe/jev-1.13 through its alpha Decisions API (POST /api/alpha/decisions) billed to your existing OpenRouter key.
Does Jev work with Cline? Not through any published integration as of October 5, 2026. jev-guard covers Claude Code, Cursor (shell/MCP), Codex, Copilot CLI, Gemini CLI, OpenCode, pi, and ACP editors — Cline isn’t on the list, so you’d wire the OpenRouter call yourself.
Is the confidence score trustworthy enough to auto-route on? The best independent evidence (60 hand-labeled tool-call risk cases, Sep 2026) found every wrong answer carried lower confidence, with Jev spreading scores more usefully (floor 0.13) than Claude (0.55+). Promising, but that’s one small benchmark — keep an “ask a human” band around your thresholds.
What does it actually cost? $0.042 per million input tokens, output free. A ~1k-token decision ≈ $0.00004; a thousand of them ≈ $0.04.
Sources
- Introducing System One Models & Jev — TypeSafe AI blog
- Models — TypeSafe AI docs
- Introducing System One Models and Jev — Hacker News discussion
- Jev on OpenRouter: Decisions API documentation
- jev-benchmark: support-triage rerun of the 193x/444x claim — arifulislamat (GitHub)
- TypeSafe’s JEV Model: Is It Really 193x Faster and 444x Cheaper? — DEV Community
- jev-benchmark: agent tool-call risk classification — themsquared (GitHub)
- jev-guard: risk-scoring hooks for coding agents — GitHub
- typesafe-cli: Jev from the shell — GitHub
- Jev and System One models — a semantic if statement for your code — dsebastien.net
- Is Jev Free? Access, Real Costs, and the Playground — Layer3 Labs
- Claude pricing reference (Sonnet 5.5 $2/$10 verified) — chudi.dev
Last updated October 5, 2026. Jev is in early access and OpenRouter’s Decisions API is alpha; pricing, limits, and availability change quickly — verify against the official pages before building on them.
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Know which coding tool is worth paying for
Hands-on comparisons of AI coding assistants and what each one costs to run — including the local-model path. Sent only when something changes. Unsubscribe anytime.