You are building an AI feature with Claude Code, Cursor, or Lovable, and somewhere in that feature there is a step that is not really "AI" at all. It is a yes/no gate. A routing decision. A risk score. A pick-one-from-a-list classification. And if you built it the way most people build AI features right now, that step is an LLM call.
That is the expensive habit worth breaking. On September 15, 2026, a new lab called TypeSafe AI came out of stealth with a model built specifically for that class of decision, called Jev. It is not a generative model and it will not write your onboarding copy. But it is a useful excuse to name a framework that will outlive Jev itself: System One calls and System Two calls belong in different places in your stack.
Key Takeaways
- Split every AI feature decision into System One calls (fast, typed, enumerable) and System Two calls (slow, generative), and route each to the right model instead of defaulting everything to a frontier LLM.
- Jev, TypeSafe AI's new "System One model," answers typed yes/no, choice, or score questions through /v1/decide, guaranteeing a schema-valid response, but not necessarily a correct one.
- At $0.042 per million input tokens with free output and 70 to 500ms latency, Jev is built to replace the habit of burning a frontier LLM call on simple classification.
- The agent-action safety gate is the clearest before/after: swapping a full LLM generation for a typed Noul call turns a per-check cost into a fraction of a cent and under half a second.
- Weigh the real caveats before relying on it: reported prompt injection risk, a single-vendor setup with no self-hosted option, and schema-valid answers that can still be wrong.
- For PMs shipping an AI feature, deciding which steps need a reasoning model, which need a fast typed decision, and which need no model at all is what makes a feature feel instant and cheap instead of laggy and expensive.
Learn this hands-on
Become a 10x PM by learning how to use Claude Code in your daily work as a Product Manager, through 3 highly efficient live sessions of 1h30. Join the Claude Code for PMs live cohort.
What Jev actually is
TypeSafe AI is a San Francisco lab founded in 2024 by Diogo Almeida, a co-author of the original InstructGPT paper. The company launched Jev with $40 million in funding (Wikipedia; The Register).
Jev calls itself a "System One model," borrowing the Kahneman term for fast, automatic, intuitive judgment as opposed to slow, deliberate reasoning. In practice, you POST your application state plus a typed question to /v1/decide on jevai.net, and it returns a typed value with calibrated confidence: a Choice (pick one option from your list), a Score (a value on a scale you define), or a Noul (yes/no with a probability attached). It cannot write code, hold a conversation, or generate prose (LangChain blog; DataCamp).
Because the response is constrained to your schema, an invalid answer is structurally impossible. That is a real property, not marketing. What it does not guarantee is that the answer is correct: TypeSafe's own CEO conceded on Hacker News that a schema-valid response can still be factually wrong. Worth keeping in your head before you gatekeep anything important with it.
Pricing is $0.042 per million input tokens, with output free, and reported latency between 70 and 500 milliseconds, a rate that puts it in the same conversation as other aggressively cheap fast models built for high-volume calls. It is currently early access, running model string Jev-1.13 (jevai.net).
The framework that actually matters: System One vs System Two calls
Jev will fade or get absorbed the way most single-purpose infra does. The framework does not have to. Split every decision your AI feature makes into two buckets.
System One calls (fast, narrow, typed): route this request to model A or B, classify this ticket into one of eight categories, score this transaction's fraud risk, gate whether this agent action is allowed to run (the same gating problem you hit once you start building agents and subagents with Claude Code), verify this output against a rubric, triage this support message by urgency. These share three traits: a fixed, enumerable answer space, a need for speed, and a volume high enough that cost compounds.
System Two calls (slow, generative, open-ended): draft the reply, explain the reasoning, write the summary, hold the conversation, generate the code. These need a frontier model because the output space is not enumerable, it is language.
Most teams building agent-driven products today skip this split entirely, which is a big part of what our harness engineering guide is about. Most teams shipping an AI feature today route both buckets through the same frontier LLM call. That is the 100x-overpaying pattern: you are asking a model that costs real money per token and takes a second or two to "think" to answer a question that has, say, six possible answers. TypeSafe positions Jev as a substitute for exactly that slot, not for Claude Code, Cursor, or Lovable, which live entirely in System Two territory (KDnuggets). It also is not really competing with traditional zero-shot classifiers so much as trying to replace the "just call a small LLM for this" habit that has become the default since GPT-mini and Haiku-class models got cheap enough to make it feel harmless.
TypeSafe's own benchmarks claim Jev is 100 to 200x faster and 200 to 445x cheaper than small frontier LLMs on classification tasks (Tom's Hardware). Treat that as a vendor claim, not an audited number, but the direction is directionally sane: a model that only ever returns one of N typed values does not need to generate a token of reasoning to get there, so it should be both cheaper and faster than one that does.
If you want the deeper cost mechanics of the LLM side of this split (when caching helps, when a bigger model is actually cheaper per useful output), that is the whole subject of our Claude API pricing and cost-performance tuning piece. And for the opposite end of the spectrum, when a decision genuinely needs maximum reasoning rather than a fast gate, see when to reach for ultracode mode.
A concrete before/after: the agent-action gate
Say you are shipping an AI feature that lets an agent take actions on a user's behalf (send an email, apply a refund, close a ticket). Before the action executes, you need a gate: is this specific action, in this specific context, safe to run automatically, or does it need a human?
Before: every gate check is a call to a small frontier model. You send it the action, the context, and a prompt asking for "safe" or "needs review," and you wait for a full generation with reasoning tokens, at $0.15 to $0.30-ish per million input tokens for a Haiku or GPT-mini-class model, plus a second or more of latency per check.
After: you define the gate as a Noul (yes/no with probability) and send the same action and context to Jev's /v1/decide endpoint. You get back a typed boolean with a confidence score in well under half a second, at a fraction of a cent per thousand checks given the $0.042/M input token rate and free output.
The delta compounds because gate checks run on every single agent action, not once per user session. If your feature runs a thousand gate checks a day, the LLM version is burning real dollars and real seconds of perceived latency on a decision that a typed classifier answers instantly. That is the whole argument in one example: nothing about the product changed, the architecture did.
This is also exactly the kind of decision our series on mastering Claude Code to ship faster and build AI agents walks through step by step: which calls belong in your agent's fast path and which belong behind a real reasoning model.
The honest caveats the hype posts skip
Jev launched to a genuinely large Hacker News reaction, over 1,500 points and well past 450 comments in a day, and Vercel and Cloudflare both shipped support within days, with community SDKs already up in Python, Go, and Elixir (Latent Space; Forbes; awesome-jev on GitHub). But the dominant read inside that thread was more sober than the launch framing: this is a very good zero-shot classifier, and calling it a "frontier model" oversells what it is.
A few things to actually weigh before you wire Jev into anything load-bearing:
- Schema-valid is not the same as correct. A
ChoiceorNoulwill always be a member of your allowed set, but the model can still pick the wrong member with high confidence. - It has reportedly been probed for prompt injection, which matters a lot more when Jev is the gate deciding whether an agent action executes, per VentureBeat's reporting. If you are using it as a security boundary rather than a routing convenience, treat that report as a reason to add a second check, not as a settled non-issue.
- It is early access and single-vendor, with no local or self-hosted option yet. That is an availability and lock-in risk for anything you would call production-critical.
- Backlash is already forming in places like r/LocalLLaMA and among the KDnuggets crowd, largely on the "this is just a classifier with better marketing" front. That critique does not make the cost and latency numbers wrong, but it is a fair check on the framing.
None of that changes the underlying lesson. Whether the tool you use is Jev, a fine-tuned BERT classifier, or a rules engine, the point stands: decide first whether a step in your pipeline is System One or System Two, and only then decide which model, if any, answers it. Jules's own rule of thumb applies here too: start with a cheap model and only escalate when the task actually needs it. If you are choosing between Claude models for the System Two half of your stack, Fable 5.1 versus Sonnet 5 as your default is the companion piece.
Shipping this in a real feature
This is exactly the kind of architecture decision that separates a demo from a feature you can actually ship: knowing which calls in your pipeline need a reasoning model, which need a fast typed decision, and which need no model at all. Getting that split right is most of what makes an AI feature feel instant instead of laggy, and cheap instead of a line item your CFO notices.
Product manager and want to work like this? This is exactly what we teach in Claude Code for PMs, our live cohort for product teams: 3 live sessions of 90 minutes over 2 weeks. Every PM ships a real feature, builds their own agent, and gets personalized written feedback.
If you would rather learn this by shipping a real product end to end, including these exact architecture calls, Ship Your First SaaS is built for that.
The bet worth making is not on Jev specifically. It is on treating "does this decision need an LLM" as a question you ask deliberately, every time, instead of a default you never revisit.