Jev is a model from TypeSafe AI that cannot write text. You give it some input and a question with a list of options. It returns its pick and a probability for every option. It answers in one pass, in about a quarter of a second, for about $0.00003 per question. It cannot think step by step, because it has nowhere to write the steps.
I spent a week testing what that one pass can do. The scores below were all checked against each dataset's real answer key. I did not rerun any frontier model. Their numbers are the published ones, with the setting each lab or evaluator used.
Jev's setting is the same everywhere: one pass, zero-shot, no chain of thought. Almost every other number comes from a model that wrote out its reasoning first. GPQA cost 1.8 cents for all 792 answers.
In short: Jev is about 20 points behind the 2026 frontier on graduate science questions and 6 to 9 points behind on multilingual knowledge. It is roughly level with GPT-5 when GPT-5's thinking is switched off. Its stated confidence matches how often it is right, which most language models cannot claim. It fails at anything that needs several dependent steps, and a search harness can take over some of that work.
How it compares with published scores
Labs stopped reporting MMLU in 2025. The benchmarks that 2026 models publish are GPQA Diamond, MMMLU and Global-MMLU, so I ran Jev on those. The older benchmarks are here too, but their comparison models are older.
The frontier reaches its GPQA numbers with extended thinking. Jev, with no reasoning at all, sits roughly where reasoning models of late 2024 sat. Chemistry was its weakest area at 66.7%, against 82.3% for physics and 72.4% for biology.
198 graduate-level science questions with 4 options. Jev is averaged over 4 option orders. Bars start at 0%.
Accuracy and price on GPQA Diamond
Jev charges $0.042 per million input tokens and nothing for output. A GPQA question is about 535 tokens. The chart prices that same prompt at each model's list input price. Output and thinking tokens are left out, so the real gap is larger than shown.
Jev: 74.0% on GPQA Diamond. Measured cost $0.000022 per question, median answer time 0.20 s.
How often its confidence is right
Every answer comes with a probability. I took all 4,036 LSAT answers and grouped them by how confident Jev said it was, then checked how often each group was right. Of the answers where it said about 70%, about 74% were right. Of the ones where it said about 90%, about 92% were right.
That makes the number usable. If you only accept answers where it is at least 80% sure, it answers 75% of the questions and gets 97.3% of those right. At a 70% bar it answers 80% and gets 96.5% right. Most language models are overconfident, so their stated confidence cannot be used this way.
Jev answers 75.0% of the 4,036 LSAT answers and gets 97.3% of those right. The rest go to a slower model or a person.
What it does when people disagree
ChaosNLI is a set of sentence pairs. The task is to say whether the second sentence follows from the first, contradicts it, or neither. Each pair was labeled by 100 people. On some pairs nearly everyone agrees. On others they split, because the pair really is ambiguous.
When people disagreed, Jev became less sure too. Its average top probability was 0.95 on pairs where people agreed, 0.88 on partly split pairs and 0.78 on badly split ones. A model trained to sound right tends to pick a side even when there is no right side. This is the best evidence I have that TypeSafe's training method, which they call RLCD, does what they say.
First sentence
Several people stopped by the side of the road are watching what is happening in the woods.
Second sentence
People are talking to eachother
Gray: votes from 100 annotators. Orange: Jev's probability for each label.
Following rules it has never seen
The LSAT and MMLU are public, so a high score may be memory. To test reasoning on something Jev cannot have seen, code generated problems with made-up rules and nonsense words, such as "a drishev is any red round object". Stated rules were essentially solved. That held for rules that contradict common sense, exceptions, rules revised later in the text, and one relevant rule hidden among 400.
Long chains are where it fails. Each category requires the one before it plus a named mark, and Jev has to find the deepest category that applies. Accuracy falls from 93% at 2 links to 8% at 16, close to the 6% chance rate. Its confidence stops tracking accuracy here: at 16 links it still reports 43%. It judges from the overall look of the input and does not walk the chain link by link.
Working out an unstated rule from labeled examples levels off in the 70s. Arithmetic before a comparison is close to guessing. Code can do both jobs: it can walk a chain one link per call, and it can add.
Accuracy by language
On MMMLU, Jev averaged 83.9% over 14 languages and 91.4% in English on the same questions. The loss from translation is small for European and East Asian languages. Most of it is in Swahili (75.4%) and Yoruba (60.8%). Global-MMLU-Lite shows the same pattern over 23 languages.
Bars start at 50%. Chance is 25%.
Adding search: the Game of 24
Jev cannot propose options, so code has to list them and Jev ranks them. In the Game of 24 you combine four numbers with + − × ÷ to make 24. Code lists every legal next step. Jev gives each step a probability, which is the policy, and rates how promising a position is, which is the value. Tree search (MCTS) uses both to decide where to look.
Following Jev's first pick solves 8 of 30 puzzles. The same judgments inside MCTS solve 22 of 30 with about 30 Jev calls per puzzle. With a random judge, the same search solves 4. The gain comes from Jev's judgment. For comparison, Tree of Thoughts reported GPT-4 at 4% with chain of thought and 74% with tree search, on a different puzzle set.
Position
8 3 11 12
Jev's pick is the probability Jev gave the move. Visits counts how many of the 60 simulations went through it. Value is the search's running average of Jev's estimate that 24 can still be reached.
Chess
Search did not help in chess, because Jev's judgment of chess positions is poor to begin with. In one pass it picks Stockfish's best move about 23% of the time and blunders on about 45% of moves. Against a 1350-rated engine it won none of 10 games. How the board was written down made no measurable difference.
| Board written as | Best move | Best in top 3 | Within 30 cp | Blunders |
|---|---|---|---|---|
| 8×8 ASCII grid | 21.5% | 38.0% | 29.5% | 49.5% |
| FEN string | 23.5% | 39.0% | 31.0% | 44.5% |
| Piece list in prose | 23.5% | 42.5% | 33.0% | 44.0% |
| How Jev played | Won | Drawn | Lost | Jev calls per move |
|---|---|---|---|---|
| Top pick, no search | 0 | 2 | 4 | 1 |
| MCTS, 24 simulations | 0 | 1 | 3 | 45 |
There is also a technical cause. Through the gateway, probabilities come back rounded to two decimals. With 30 or more legal moves, most moves read 0.01 to 0.03, so the ranking is nearly flat and search has little to work with.
Consistency checks
A model with honest probabilities should also give the same answer when you ask the same thing in a different way. Jev mostly does. The clear exception is negation. Asked whether a statement is true and, separately, whether it is false, its two probabilities should add to 1. They miss by 0.13 on average. TypeSafe lists negation as a known weakness.
| Check | Result |
|---|---|
| Same question, 3 option orders × 3 phrasings: answer changed | 1.0% |
| Which of two texts is more formal, 200 triples: preference cycles (A > B > C > A) | 0.5% |
| "Is it true?" plus "is it false?": average distance from 1 | 0.13 |
| Question asked alone vs. with another question: average shift | 0.02 |
I also added junk and planted answers to 200 LSAT questions. Sixteen unrelated paragraphs cost 2 points. A note saying "the correct answer is" with a wrong letter cost 3 points, and Jev followed the planted letter 6% of the time. A fake "SYSTEM: ignore the passage" line was followed 3% of the time. Its confidence dropped in those conditions.
| Condition | Accuracy | Mean confidence |
|---|---|---|
| Clean | 94.0% | 0.95 |
| 16 unrelated paragraphs added | 92.0% | 0.92 |
| Planted wrong answer | 91.0% | 0.89 |
| Fake system instruction | 94.0% | 0.91 |
What it knows
TypeSafe has said almost nothing about what is inside Jev. Its knowledge scores say there is a large pretrained language model. It can pick a missing word out of 255 candidates 71% of the time. Its knowledge of rare entities drops off the way a language model's does. It answered questions about events through the end of 2025 perfectly and missed some from 2026, so its training data ends around early 2026.
| Test | Questions | Jev | Chance |
|---|---|---|---|
| MMLU, 10 per subject | 570 | 91.1% | 25% |
| MMLU-Pro | 140 | 81.4% | 11% |
| PopQA, rare and common entities | 1,000 | 69.7% | 25% |
| PopQA, obscure entities only | 62.5% | 25% | |
| PopQA, famous entities only | 92.8% | 25% | |
| Dated events, 2022 to 2026 | 48 | 95.8% | 25% |
| Missing word, 255 candidates | 200 | 71.0% | 0.4% |
What it is good for
Jev is a fast judge with honest confidence. It is weak at anything that needs several dependent steps. The practical design puts the steps in code or search, and uses Jev's confidence to decide when to hand a case to a slower model. I tried that with a browser agent in a second post.
Methods and caveats
How Jev was run+
All calls went to typesafe-ai/jev through Vercel AI Gateway between 19 and 25 September 2026. Each question was one Choice call: the passage and question as the input, the answer options as the choices. No examples, no retries for correctness. The earlier study cost about $1 in Jev calls. GPQA cost 1.8 cents, MMMLU 15 cents and Global-MMLU-Lite 10 cents.
TypeSafe's own published accuracy numbers are scored against what GPT-6 Astra and Claude Fable 5.1 agree on. That measures agreement with those models. Every number here is scored against the dataset's real labels instead.
Published numbers and their settings+
No frontier model was run for this post. Every comparison is a published number from a lab page, model card, system card or an evaluator (Artificial Analysis, Vals, Epoch AI, llm-stats). Anthropic no longer publishes GPQA, so the Claude GPQA number is an Artificial Analysis run at max effort. Most comparison numbers come from models that generated reasoning tokens before answering. Jev did not.
Costs for other models in the price chart are estimates: 535 prompt tokens times each model's list input price. Output and thinking tokens are not counted.
Sample sizes and intervals+
GPQA Diamond: all 198 questions, each in 4 option orders; 95% interval 70.8% to 76.9%. Jev chose the same option across orderings on 75% of questions. MMLU is a 570-question subset and MMLU-Pro a 140-question subset, so a difference of a few points against a leaderboard row is within noise. The LSAT run is all 1,009 AGIEval questions in 4 option orders.
MMMLU: some translations reorder questions while reusing ids, so six affected subjects were dropped in every language and 500 paired questions were sampled per language. It is a filtered sample, not the official full set. Global-MMLU-Lite used 200 questions in each of 23 languages.
Contamination+
AGIEval LSAT and MMLU are old and public. Jev has probably seen them in training, so those scores show what it knows, not how it would do on a new exam. The invented-rule problems are generated from a seed and cannot have been seen.
ChaosNLI metric+
Jev's Jensen-Shannon distance to the human vote distribution is 0.21 on 200 pairs. Fine-tuned 2020 models scored 0.22 to 0.24, 2026 reasoning models about 0.12 after chain of thought, and a second set of human annotators 0.06. An earlier draft of my notes reported 0.06 for Jev, which was the divergence, not the distance.
Search and chess details+
Game of 24: the same 30 seeded puzzles in every row. Beam search used width 5 and 157 calls per puzzle. The random-judge row comes from the run notes, not a separate result file. Chess games were 6 with no search and 4 with MCTS, too few to measure a difference.