MLGuerrillaStart with M1 →
Text and language·16 min read·Updated 23 September 2026

Jev and System One models

Jev is a model from TypeSafe AI that reads text and answers questions whose answers you list in advance. It returns a probability for every answer you listed, and it never writes text.

Jev is a model from TypeSafe AI, released in early access on September 15, 2026. It takes text as input, like a support ticket, together with questions you write about that text. Each question is written in advance with the full list of answers it's allowed to have. For example, "Which team should handle this ticket?" could allow only billing, technical and sales as answers. Jev returns a probability for each allowed answer, and your code decides what to do with them. It can't write a sentence, so it can't reply to a customer or explain itself.

TypeSafe calls this kind of model a System One model. The name comes from Daniel Kahneman's book Thinking, Fast and Slow, where System 1 is the fast, intuitive kind of thinking and System 2 is the slow, deliberate kind. As of September 18, 2026, Jev is the only System One model TypeSafe has released, so this page covers both the model and the idea behind it.

What it is

Jev answers questions your code defines, with a probability for every answer

Most of the AI inside ordinary software is making small decisions. Which team should get this ticket, and is the customer angry enough that somebody should call them back? The usual way to answer a question like that today is to send the text to an LLM and ask it to reply in JSON. Your code then has to parse what comes back and check that it matches the shape you asked for, because the model can write anything.

Jev takes two inputs:

  • The state is the content you want judged. It can be plain text, and it can be JSON when the decision needs several pieces together, like a customer's message alongside your refund policy.
  • The questions each carry a type and the full list of answers they are allowed to return, filed under an ID you choose so your code can find the answer again.

There are three question types, and each returns a different shape of answer.

The three question types (TypeSafe docs, as of September 18, 2026)
TypeWhat it asksWhat comes back
ChoiceWhich of these options fits? Up to 255 options, each with a description you write.choice (the most likely option), probabilities for every option, and confidence
ScoreWhere does this sit on a scale? You describe 2 to 10 ordered levels.score, the probability-weighted average of the level numbers, plus probabilities and confidence
NoulIs this statement true?noul, the probability that the answer is yes

TypeSafe's quickstart sends one support message, "Hi, I've been trying to connect my Stripe account for 3 days and it keeps failing. I'm losing sales. Please help ASAP.", with three questions. The department question is a Choice between billing, technical and sales. The frustration question is a Score with three levels, from "Calm, just stating facts" (0) to "Very angry, strong language" (2). The is_urgent question is a Noul. This is the answer from the docs, with the Score's legend field left out:

Jev's response in TypeSafe's quickstart (trimmed)
{
  "model": "jev-latest",
  "answers": {
    "department": {
      "type": "choice",
      "choice": "billing",
      "probabilities": {"billing": 0.84, "technical": 0.159, "sales": 0.001},
      "confidence": 0.596
    },
    "frustration": {"type": "score", "score": 1.035, "confidence": 0.842},
    "is_urgent": {"type": "noul", "noul": 0.999}
  },
  "usage": {"input_tokens": 312, "output_tokens": 48}
}

Billing wins with 0.84, but technical still gets 0.159, because a Stripe connection that keeps failing could be a payment problem or an integration bug. The frustration score of 1.035 sits just past level 1, "Frustrated but civil". The message is almost certainly urgent at 0.999. Your code branches on these numbers the way it would branch on any other value, which is why TypeSafe's launch post describes these calls as "smart if-statements".

Where it shows up

It fits the small decisions inside software that run thousands of times a day

The jobs TypeSafe built Jev for all look the same from the outside. You know every possible answer before you make the call, the same question gets asked constantly, and code acts on the answer without a person ever reading it.

Routing and triage

A support system asks several questions about each ticket in one call. The department is a Choice, and questions you might not need go in the same call, like the return reason, which is only useful when the ticket turns out to be about a return. The code routes the ticket to the top team and sends a copy to any other team with a probability above 0.25. If the question about what the customer wants comes back with low confidence, the code asks the customer before acting. TypeSafe's Choice docs walk through this exact example, and the only logic in it is ordinary if statements.

Deciding which bigger model to call

Jev can sit in front of more expensive handlers and decide which one each request needs. In TypeSafe's intent-routing pattern, order-status questions go to plain code with a database lookup, product questions go to an LLM loaded with product information, and complaints that score as complex go to a person. The expensive LLM only runs on the requests that need it.

Checking what another model wrote

A Noul like "Does this draft reply promise the customer a refund?" can run on every LLM reply before it's sent. TypeSafe lists judging LLM output and detecting jailbreaks among its intended uses. Text inside the state can steer Jev's answer, though, which is a problem for a guardrail in particular (the limits section covers it).

Turning a large pile of text into columns

Running the same questions over millions of records turns free text into features you can chart or feed to a classical model. Every product review could get a complaint category and a sentiment score, for example. At $0.042 per million input tokens, a billion tokens of reviews costs $42.

Real-time loops

Because a call takes well under a second by TypeSafe's numbers, Jev can make decisions inside a loop that can't wait for an LLM. TypeSafe's launch demo plays Doom by asking Jev about the game 10 times a second, which TypeSafe says costs about $7 an hour. The game state goes in as a data structure with text in it, because Jev can't read images.

How it works

Jev reads the state once and answers every question in parallel

How an LLM would answer the same question

An LLM is a transformer, the architecture from Attention Is All You Need (Vaswani et al., 2017), trained to predict the next token. Tokens are the chunks of text a model reads and writes, often a word or part of a word. To answer "which department?", the LLM writes its answer one token at a time, and each new token is fed back in before it picks the next one. This loop is called autoregressive decoding, and it's why generation gets slower as the output gets longer.

If you ask the LLM for a probability too, the characters "0.84" are generated the same way as the word "billing". They're the text the model found likely to come next, and nothing in the loop checks them against how often the model is right. Xiong et al. (2023) tested confidence that LLMs state in words and found that they "tend to be overconfident". The GPT-4 Technical Report showed a related effect in the model's own token probabilities. On multiple-choice MMLU questions, the pre-trained model's probabilities matched how often it was right, and the report's Figure 8 says "post-training hurts calibration significantly".

TypeSafe publishes an open-source adapter that gets Jev's output format out of OpenAI and Anthropic models, and its code shows how much work that takes. The system prompt asks the model to "Preserve genuine uncertainty" and to "make the probabilities sum to 1". Then the code rescales any distribution that doesn't add up to 1, and resends the request with a correction message when the JSON doesn't match the schema.

What Jev does with the same request

TypeSafe's launch post says Jev outputs all of its probabilities in parallel in a single query, with no token-by-token generation. The docs add the rest of the published picture:

  • Jev reads the state once and evaluates every question against it in parallel.
  • Each question is evaluated in isolation, so one question's answer is never context for another.
  • The answer to each question is a probability distribution over the options you listed. An answer outside your list can't come back, because there's no generated text for it to come from. The same rule means a ticket that fits none of your options still gets one of them, so the docs say to add an other or none of the above option when the list might not cover every input.

That last point is what TypeSafe means when it says type errors are "mathematically impossible" and that Jev "can't hallucinate". A wrong answer is still possible, since Jev can put most of the probability on the wrong option. So the guarantee is about the shape of the answer, and only your own evaluation can tell you how often the answer is right.

Parallel questions also change how you design a request. Adding a question barely changes the response time, and the only extra cost is the question's own tokens. TypeSafe's parallel-questions cookbook reports that batching 13 questions into one call "is 11.5x cheaper and 9.6x faster than 13 separate calls". So the docs tell you to ask every question you might need in the same call, then have your code ignore the answers it doesn't use.

A figure titled "Jev reads the ticket once and answers every question in parallel", with two columns. The left column, "An LLM writes its answer as text", starts from a support ticket card quoting "…my Stripe account… keeps failing… help ASAP." An arrow leads to a dark box, "LLM predicts one next token, then reads it back in", and then to a chain of numbered token chips in two rows. The chips are an opening brace, "department", a colon, "billing", a comma, "p", a colon, 0.84 and a closing brace, numbered 1 to 9, followed by a faded chip with an ellipsis. A note says one step per token, and that frustration and is_urgent would add more steps to the same chain. The chain ends at a card reading "Your code parses the JSON, retries if it's broken". The right column, "Jev scores the options you listed", starts from the same ticket. An arrow leads to a dark box, "State read once and shared by every question", and three terracotta arrows fan out from it to three cards side by side. The first, a Choice named department, shows bars for billing 0.84, technical 0.159 and sales 0.001. The second, a Score named frustration, shows a 0 to 2 scale with a marker at 1.035, near level 1, frustrated but civil. The third, a Noul named is_urgent, shows a bar for yes at 0.999, the probability of yes. A note under the cards says each question sees the state and its own options only, and the column ends at a card reading "Your code branches on the numbers, with nothing to parse". The caption at the bottom says that Jev's probabilities are computed over the options you listed, so an answer outside your list can't come back and there's no JSON for your code to repair.
Adding a question on the Jev side adds one more card to the same pass, which is why TypeSafe says extra questions barely change the response time.

What's inside is unpublished

TypeSafe describes "a new model architecture" and a "parallel sampler", and its FAQ says Jev "is neither small nor an LLM". As of September 18, 2026, it hasn't published a paper, and it hasn't said how many parameters Jev has or what base model it started from. It also says it makes all of its training data itself and doesn't train on customer requests.

The most detailed outside attempt to work out the architecture is an essay on archerhume.com, based on about 10,000 calls to jev-1.13.0 made on September 17, 2026. Its author argues that Jev is most likely a transformer that stops once it has read the input and never enters the decoding loop at all:

  • The state is processed once and cached, and every question runs as its own branch that can read the state but can't read the other questions. When the author put a secret code inside one question, a second question couldn't find it (probability 0.00). With the same code moved into the state, the second question found it with a probability of 0.90 to 0.92.
  • The options inside one question are read together. Adding an irrelevant fifth option, "Bad weather caused it", to four causes of a payout failure changed the odds between two of the original options. Scoring each option on its own would have left those odds unchanged.
  • Server time barely changed as a request grew from one question to about 100, which fits the state being computed once and the questions being batched.
  • A small output layer turns each branch's final internal representation into one score per option, and softmax, the function that turns scores into probabilities that sum to 1, gives the distribution you see.

The author guesses the backbone is a mixture of experts (a model that runs only part of its weights for each token) and marks that as the least certain part. All of this is one person's inference from the outside, and TypeSafe hasn't confirmed any of it.

If the reconstruction is right, the output side is doing the same thing a fine-tuned classifier does. A BERT-style encoder reads the whole input in one pass, and adding one output layer lets you fine-tune it to put a probability on each of your labels. The difference is that a BERT classifier learns a fixed list of labels from your training examples, while Jev takes a new list of options in every request, described in plain words, with no training on your side.

TypeSafe says RLCD trains the probabilities to match outcomes

A model is calibrated when its probabilities match how often it's right. Across all the answers it gives 0.8, about 80% should turn out correct, and across the answers it gives 0.2, about 20%. Calibration is measured over many predictions, so it says nothing about whether any one answer is right. Guo et al. (2017) showed that modern neural networks are often poorly calibrated out of the box, and that a small correction after training, called temperature scaling, fixes much of it.

TypeSafe's launch post says "All answers are accompanied with calibrated probabilities and confidence scores", and it publishes no calibration data to back that up, with no reliability chart and no calibration error number. Measuring calibration only needs the model's answers and the correct labels, so anyone with API access can check it by running Jev on labeled test cases and comparing its probabilities with how often it was right. The how-to section below describes the chart. Jev is still in early access, so the measurement below is on a model anyone can run. The classifier is the fine-tuned encoder from the BERT article, scored on all 3,080 held-out messages.

A figure titled "A calibrated model is right about as often as it says it will be", noting that Jev is in early access, so this is a classifier anyone can run, measured on 3,080 held-out messages. On the left is a plot with what the model said along the bottom, from 0% to 100%, and accuracy up the side. A dashed diagonal marks perfect calibration. Nine terracotta dots, one per 10-point band and sized by how many answers landed in it, sit close to the diagonal at the high end and well below it at the low end. The largest dot holds 2,696 answers near 99% confidence and 98% accuracy. A band of 22 answers sits at about 36% confidence and 23% accuracy, and a band of 13 answers sits at about 27% confidence and 8% accuracy. On the right, three figures over 3,080 answers: it said 95.1% on average, it was right 93.4% of the time, and the average gap between the two is 1.8 points. Below that, one band is read out loud, saying that on 22 answers the model gave about 36% and was right on 5 of them, which is 23%, and that a threshold set at 0.3 would let through mostly wrong answers. The caption reads: the number a model hands you is only useful if it matches how often that number turns out to be right. A note under it says Jev's probabilities would be measured the same way.
A gap of 1.8 points across all answers hides a much wider gap in the low bands, which is where a threshold actually gets used.

Two things in that measurement carry over to any model that hands you a probability. The average looks good while individual bands are far off, so an overall number hides the bands where your thresholds do their work. And almost every answer lands in the top band, which means a few hundred test examples tell you almost nothing about the bands below it.

TypeSafe's primer places its training method next to the two post-training methods that shaped today's LLMs:

  • RLHF, reinforcement learning from human feedback, trains a model to write the responses people prefer. It's the method behind InstructGPT and ChatGPT, described in Ouyang et al. (2022). TypeSafe's founder, Diogo Almeida, is one of that paper's authors.
  • RLVR, reinforcement learning with verifiable rewards, trains on answers a program can check, like math results. The name comes from the Tülu 3 paper (Lambert et al., 2024), and it's the kind of training behind today's reasoning models.
  • RLCD, reinforcement learning for calibrated decisions, is TypeSafe's own. TypeSafe says it trains the model to return decisions whose probabilities match outcomes.

TypeSafe's argument is that RLHF rewards answers that sound good to a person, which can include confident-sounding mistakes, so it's the wrong objective for decisions nobody reads. The RLCD recipe isn't published. The standard way to train for calibration is a proper scoring rule, like log loss or the Brier score, where the model gets the best score only by reporting the probabilities it believes. TypeSafe hasn't said which one it uses, or how it gets the correct outcomes to train against.

The confidence field on Choice and Score answers summarizes how peaked the distribution is, on a 0 to 1 scale. It's computed from the probabilities, so it carries whatever calibration they have and no more. In the quickstart response, billing's 0.84 with 0.159 left on technical gives a confidence of 0.596, lower than the 0.84 on its own, because the runner-up still holds real probability. Noul answers have no separate confidence, since the probability of yes already says how sure Jev is.

Versions and pricing

The version you'll see is jev-1.13, through TypeSafe's API only

Jev as of September 18, 2026 (docs.typesafe.ai)
Current value
Model IDjev-1.13.0, the only released version
Aliasesjev-latest and jev-preview, both pointing to jev-1.13.0
Price$0.042 per million input tokens ($42 per billion). Output tokens are free.
Context64k tokens for the state plus all questions, and 32k for the state plus the longest question
Rate limits250,000 tokens per second and 1,200 requests per minute, which TypeSafe says can change without notice
InputText only, as a string or as JSON. English gives the best accuracy.
WeightsClosed. Every account uses the same weights, and there's no fine-tuning.
AccessEarly access from a waitlist since September 15, 2026, with a Python and a JavaScript SDK

An alias moves to the new model when a new version ships, so the numbers you get back can change with no change on your side. The response's model field reports the exact version that answered. If you've tuned confidence thresholds against one version, pin its ID and move to the next one when you've re-run your evaluation.

Choosing

The shape of the answer decides whether Jev or an LLM fits the job

Two questions settle most choices. Can you write down every possible answer before the call? And do you have the labeled data and the reason to train and run a model of your own?

Pick Jev when you can write the possible answers down

  • The answer is one of a list, a level on a scale or a yes/no, and you know the list before the call.
  • You make the decision thousands of times a day, or inside a loop that needs an answer in well under a second.
  • Your code needs a probability it can threshold, like sending low-confidence tickets to a person, and you've checked on your own labeled tickets that the probabilities hold up.
  • The labels change often. Adding a department means adding an option to the request, with no retraining.
  • You don't have labeled data to train a classifier.

Pick an LLM when the output is text or the answer needs reasoning

  • The job is to write something, like a reply or a summary.
  • The answer is a value you can't list in advance, like a name or an amount pulled from an email. TypeSafe suggests pulling out candidate values with a regular expression or an LLM first, then asking Jev to choose between them.
  • The decision needs several steps of reasoning. TypeSafe's own docs also say to keep arithmetic and date comparisons in your code.
  • The input is an image or audio, since Jev only reads text.
  • The volume is low, so a few seconds and a few cents per call are fine.

Pick a classifier you train when the labels are fixed and the data can't leave

A fine-tuned encoder, or text embeddings fed into a logistic regression, fits a different situation. You have thousands of labeled examples and labels that rarely change. You need to run the model on your own hardware because the data can't go to a third-party API, or because you need a model that stays available for as long as you want it. Jev runs only on TypeSafe's API, so every state goes to TypeSafe. Pinning jev-1.13.0 keeps the version fixed while TypeSafe serves it, and TypeSafe decides how long that is.

A worked example on cost

Take an invented company that triages 200,000 support tickets a day, with about 500 input tokens per ticket once the message and the questions are included, and three questions per ticket. For the LLM, assume it writes about 80 tokens of JSON per ticket. The volume and token counts are made up, and the prices are list prices as of September 18, 2026.

Cost of 200,000 ticket decisions a day (invented volume, list prices as of September 18, 2026)
Input tokens a dayOutput tokens a dayCost a day
Jev (jev-1.13.0)100Mnot billed$4.20
Claude Haiku 4.5, standard API100M16M$180
Claude Haiku 4.5, batch API (results in up to 24 hours)100M16M$90

Jev's line is 100 million tokens at $0.042 per million. Haiku 4.5's standard line is $100 of input at $1 per million and $80 of output at $5 per million, and Anthropic's batch API halves both prices. The real LLM bill would move in both directions, because an LLM prompt usually carries a longer system prompt and a JSON schema, while prompt caching makes the repeated part of the prompt cheaper. Different models also count tokens differently. The table also leaves out accuracy, which you'd need to measure on your own tickets before the cost is worth comparing.

What the speed and quality claims rest on

TypeSafe says a call takes 70 to 500 milliseconds, against 3 to 329 seconds end to end for frontier LLMs, and calls Jev 40 to 200 times faster "for the same levels of frontier intelligence". Its workflow evaluations put Jev at 193.6 times faster and 444.6 times cheaper, and the launch post says it expects those to be "on the higher end of real world gains". The launch post also lists the caveats itself:

  • The reference answers are the average of GPT-6 Astra and Claude Fable 5.1, so the scores measure how closely each model agrees with those two, and no labeled ground truth is involved.
  • TypeSafe's own team wrote the workflows, and TypeSafe says "some bias could exist".
  • The LLMs ran through TypeSafe's wrapper, which TypeSafe says tends to be slower and more expensive than asking an LLM for a plain answer with no probabilities.
  • TypeSafe chose not to publish results on any public benchmark.

So treat every speed and quality number in this section as TypeSafe's own until someone measures them independently.

Try it

How to try it

You need early access first, which comes from the waitlist on typesafe.ai. After that there are three ways in:

  • The Playground at console.typesafe.ai, where you paste a state and questions and see the answers.
  • The HTTP API, a POST to https://api.typesafe.ai/v1/systemone with your key.
  • The Python SDK (pip install typesafe-sdk, Python 3.10 or newer) or the JavaScript SDK (@typesafe-ai/sdk).

This is the quickstart call in Python, cut down to two questions. The client reads TYPESAFE_API_KEY from the environment and uses jev-latest unless you name a version.

triage.py — two questions about one ticket
from typesafe_sdk import Choice, Noul, TypeSafeClient

client = TypeSafeClient()
ticket = "My Stripe account keeps failing to connect. I'm losing sales. Please help ASAP."

response = client.system_one(
    state=ticket,
    questions={
        "department": Choice(
            instructions="Which team should handle this",
            criteria={
                "billing": "Payment or subscription issues",
                "technical": "Bugs or integration problems",
                "sales": "Pricing or account questions",
            },
        ),
        "is_urgent": Noul(instructions="The message conveys urgency or time-sensitivity"),
    },
)

department = response.answers["department"]
print(department.choice, department.probabilities, department.confidence)
print(response.answers["is_urgent"].noul)

To decide between Jev and an LLM for your own job, run both on the same labeled test cases. TypeSafe's open-source system-one-adapter (MIT license) exposes the same system_one call backed by OpenAI or Anthropic models, so one script can send identical questions to both. Datasets & Evaluation covers building the labeled dataset. For calibration, group the answers by their top probability, for example 0.9 to 1.0, 0.8 to 0.9 and so on, and check whether the share that's correct in each group matches the probability. Guo et al. call this chart a reliability diagram.

Ask your AI coding tool

Build a small evaluation that compares Jev with an LLM on support-ticket routing. Read tickets.csv, which has a text column and a human-labeled department column (billing, technical or sales). For each ticket, ask Jev one Choice question with those three options using the typesafe-sdk Python package, and ask the same question through system-one-adapter with llm_answer_mode="probabilities" and a model I'll name. Record the full answer and the call's latency. Report accuracy and latency (median and 95th percentile) for each model. Then add a calibration table that groups answers by top probability in 0.1-wide bins and shows the accuracy in each bin. Keep the API keys in environment variables and don't hardcode any thresholds.

Limits

What it can't do

TypeSafe publishes its own list of Jev's failure modes for jev-1.13, last reviewed on September 17, 2026, and says many of them will be fixed in later versions.

  • It can't write text. You can force it to spell something out by chaining Choice questions, and TypeSafe says that works badly and slowly.
  • It can't count or do arithmetic reliably. That includes counting items in a list and reading an exact number off a Score. Do the math in code, or ask one question per item and add up the answers yourself.
  • It can't compare dates reliably. Have it pull out the parts of each date as Choice questions, then compare the dates in code.
  • It can't answer "none of these" unless you list it. A Choice spreads all of its probability across the options you wrote, so a spam message sent to a triage question with billing, technical and sales still comes back as one of the three. Add an other option with its own description, which TypeSafe's docs recommend, and route that answer to a person.
  • It reads instructions literally. It answers the words you wrote, including negations, so when you find yourself explaining what you meant, that explanation belongs in the question.
  • Questions with double negatives or several hops of reasoning get less reliable answers. Point the question at the exact field in the state, and split a multi-step judgment into separate questions.
  • Accuracy drops when the state is full of material the question doesn't need. Filter in code first and send only the relevant fields.
  • Text in the state can steer the answer. Jev doesn't treat the state as hostile, so an injected instruction or a message that argues for its own classification can move the probabilities. That's a problem in particular when Jev is the guardrail.
  • Related questions don't have to agree. On one ticket in TypeSafe's docs, "Is the customer asking for a refund?" came back at 0.72 and its negation at 0.47, which adds up to 1.19. A threshold you tuned on a Noul doesn't carry over to a Choice asking the same thing.
  • It only reads text, and English works best. Other languages work with lower accuracy, and anything that isn't text has to be converted to text before it goes in.
  • Nothing about it is open. The architecture and the weights are unpublished, and it runs only on TypeSafe's API, so you can't self-host it or fine-tune it on your own labels.
  • Calibration is a property of many answers. If Jev is calibrated, a probability of 0.9 means about 90% of such answers are right across a large group. TypeSafe hasn't published that measurement, so your thresholds have to come from testing on your own data.

Coming soon

New modules go live as I write them. Get each one in your inbox the day it ships. No spam, just the next lesson.