“System One” is TypeSafe‘s name for the model class, and it’s borrowed from psychology: the phrase comes from Daniel Kahneman’s Thinking, Fast and Slow. System 1 is the quick, automatic judgment you make before you finish reading a sentence; System 2 is slow, careful working-out. TypeSafe’s founding argument is that frontier labs are racing to build a better System 2, but most decisions inside real software are System 1 questions. And we’ve been paying System 2 prices to answer them. “Is this urgent?”, “which queue?”, “spam or not?” these kinds of judgment your triage webhook makes, they don’t need an essay-writing engine behind it.
The simplest way to think about it is:

It is built for decisions, not language
A normal LLM is optimized to produce strings that humans find useful: explanations, code, emails, conversations and so on.
Jev is deliberately giving up that flexibility. You ask something like:
State:
Customer has contacted support twice.
Account has been active for 4 years.
Customer says they may cancel.
No prior retention offer.Question:
How likely is churn?Allowed output:
LOW
MEDIUM
HIGH
The output is conceptually more like:
HIGH: 0.78
MEDIUM: 0.18
LOW: 0.04
rather than: “Based on the customer’s long tenure but recent dissatisfaction, I would classify them as…”
That makes the output directly usable by code. TypeSafe describes this as “typed probabilistic decisions” rather than text.
The output space is defined in advance
This is probably the biggest architectural difference. With a generative LLM, theoretically: output = almost any string. Even when you ask for JSON, the underlying model is still generating tokens.
TypeSafe currently exposes three main question types:
- Noul: roughly a yes/no probabilistic judgment
- Choice: select among defined alternatives
- Score: place something on an ordered scale
Their public API exposes these as native schema types. For example:
- Noul: “Is this transaction suspicious?”
- Choice: “What is the support intent?”
- refund
- billing
- technical
- cancellation
- Score: “How severe is this incident?”
- 1 → 5
This means the surrounding software already knows what kinds of answers can appear.
It doesn’t sample one token at a time
Normal LLM inference is autoregressive:
token 1
↓
token 2 conditioned on token 1
↓
token 3 conditioned on tokens 1+2
↓
…
That sequential generation is one reason output can be expensive and relatively slow.
TypeSafe says Jev uses a parallel sampler instead. It produces its structured outputs together rather than generating a long textual sequence one token at a time. This is a major reason they claim substantially lower latency and cost on System-One-shaped workloads.
That makes it more like:
Input state
↓
model inference
↓
↓ ….↓ ….↓
Q1 …Q2 …Q3
0.9 …B …High
rather than three paragraphs of generated reasoning.
Confidence is part of the interface
This is one of the most interesting parts for automation. A conventional LLM might say: “I’m 90% confident…” But this number is generated text and may not be well calibrated. TypeSafe’s stated goal is for probability/confidence to be an intrinsic part of the output. Their training approach is called Reinforcement Learning for Calibrated Decisions (RLCD).
The objective is that, over time: decisions emitted with ~90% confidence, should be correct roughly ~90% of the time. This matters because software can then make policy decisions around uncertainty. For example:
confidence ≥ 0.95
↓
automatic action0.70–0.95
↓
additional checks< 0.70
↓
human review
That is much more useful operationally than simply getting a classification without knowing how uncertain the model is.
The workflow does the reasoning structure
This is perhaps the most important conceptual point. System One is not saying: “Put the whole business problem into one prompt and let Jev solve everything.” Instead, it pushes you toward decomposing the problem into a workflow. For example, invoice processing could look like:
Invoice arrives
↓
Is this actually an invoice?
↓
Is supplier valid?
↓
Fraud indicators?
↓
Duplicate payment?
↓
Amount discrepancy?
↓
Approval required?
↓
CODE
↓
Pay / Hold / Escalate
Some steps are ordinary deterministic code. Some steps require fuzzy judgment and go to Jev. TypeSafe’s own workflow evaluations emphasize exactly this pattern: decompose the task into narrow judgments and leave arithmetic, dates, account numbers, rules and other deterministic operations to code.
So System One architecture is closer to: Deterministic logic + small AI decisions + deterministic logic + small AI decisions -> business outcome
rather than: Give everything to LLM -> “figure it out”
That’s a substantial philosophical difference.
Why they say “type-safe”
Suppose your application expects exactly one of three values: “approve” | “reject” | “review”. Your code branches on this value. Anything else breaks it.
Now ask a generative model for the decision. An LLM writes its answer as free text, one token at a time. Free text can be anything, so alongside the answer you wanted, it can theoretically produce:
- “APPROVE” – right idea, wrong casing, fails a string match
- “probably approve” – hedging, not one of your three values
- “approve_pending” – a value it invented
- “I cannot determine…“- a paragraph instead of an answer
Structured-output features greatly reduce this risk, but they can’t remove it. The model is still writing the answer, and anything written can go off-script.
Jev doesn’t write its answer. You define the three options up front, and the model’s only job is to distribute probability across them. It’s a multiple-choice sheet, not an essay: a fourth option is structurally impossible. The output belongs to your type by construction. That’s why TypeSafe can make the strong claim that schema and type errors are impossible for Jev’s output interface.
But there’s an important distinction hiding here: Type-correct does not mean decision-correct.
Jev will always return one of: “approve” | “reject” | “review”, but it can still pick the wrong one of the three. A perfectly parseable “approve” on a loan that deserved “reject” passes every type check and is still a bad decision. We need to distinguish “no hallucination” from “no mistake”.
That’s why calibration and workflow design remain important.
System One vs System Two is a useful architecture pattern
The really interesting possibility is not replacing frontier LLMs. It is splitting AI workload through workflow design. Think of:
System One: Fast, repetitive, bounded decisions:
- classify
- route
- score
- rank
- detect
- filter
- verify
- choose
System Two: Slow, open-ended reasoning
- investigate
- plan
- explain
- synthesise
- write
- reason through ambiguity
- develop strategy
Then a system could look like:

That can potentially give you better economics than using a large reasoning model for every tiny decision.
Where this becomes powerful for AI agents
Consider an AI agent dealing with ITSM incidents. Today you might give the whole incident to an LLM:
- Read incident.
- Determine severity.
- Determine responsible team.
- Decide whether this is security-related.
- Choose the correct runbook.
- Decide whether human approval is required.
- Then execute.
System One encourages decomposition:
Incident
│
┌────────┴─────────┐
▼ ▼
severity score security issue?
│ │
└────────┬─────────┘
▼
deterministic policy logic
│
┌────────┴─────────┐
▼ ▼
known runbook? ambiguity high?
│ │
▼ ▼
execute safely frontier LLM /human analyst
This workflow-reset resembles traditional enterprise systems much more closely. AI becomes an intelligent component inside controlled software, rather than the software surrendering control to an LLM. Remember what we discussed in Stop Fitting AI Into People-Shaped Processes?
The deeper idea behind System One
To me, the most important idea isn’t Jev itself. It’s this architectural shift: we don’t necessarily need AI to generate an explanation every time software needs intelligence.
For the last several years we’ve largely treated LLMs as the universal primitive: problem → prompt → text
System One proposes another primitive: state → bounded question → probabilistic decision
And once you have that primitive, you can build much more conventional software around AI. For enterprise automation and especially governed agentic AI, that is a very interesting direction.
The caveat is that TypeSafe/Jev is still extremely new, and most of the published performance evidence currently comes from TypeSafe itself. Their workflow evaluations are useful, but I’d regard claims such as 193× faster or 444× cheaper as vendor benchmark results until independently reproduced across broader workloads.