Skip to content
← Blog
AIEvals

Testing Jev

Jev is a model from TypeSafe AI that cannot write text. You give it some input and a question with a list of options. It returns its pick and a probability for every option. It answers in one pass, in about a quarter of a second, for about $0.00003 per question. It cannot think step by step, because it has nowhere to write the steps.

I spent a week testing what that one pass can do. The scores below were all checked against each dataset's real answer key. I did not rerun any frontier model. Their numbers are the published ones, with the setting each lab or evaluator used.

Scorecard
GPQA Diamond74.0%Best published: GPT-6 Astra 96.0%, thinkingFor reference: Human PhD experts about 65%. GPT-4 (2023) about 36%.
MMMLU, 14 languages83.9%Best published: Gemini 3.1 Pro 92.6%, thinking highFor reference: GPT-5 with thinking off 84.9%. GPT-4o 81.4%.
Global-MMLU-Lite86.0%Best published: Gemini 3.1 Pro 93.2%. Claude Opus 5.5 94.3% on the full 42-language set, max effort.For reference: Gemini 3 Pro 92.2%.
MMLU91.0%Best published: GPT-5 92.5%, self-reported, setting not statedFor reference: GPT-4 86.4%, 5-shot.
MMLU-Pro81.4%Best published: Gemini 3 Pro 89.8%, thinking highFor reference: Llama 4 Maverick 80.9%, no reasoning.
LSAT (AGIEval)86.9%Best published: GPT-4 74.4%, few-shot, 2023For reference: Human average 56%.

Jev's setting is the same everywhere: one pass, zero-shot, no chain of thought. Almost every other number comes from a model that wrote out its reasoning first. GPQA cost 1.8 cents for all 792 answers.

In short: Jev is about 20 points behind the 2026 frontier on graduate science questions and 6 to 9 points behind on multilingual knowledge. It is roughly level with GPT-5 when GPT-5's thinking is switched off. Its stated confidence matches how often it is right, which most language models cannot claim. It fails at anything that needs several dependent steps, and a search harness can take over some of that work.

How it compares with published scores

Labs stopped reporting MMLU in 2025. The benchmarks that 2026 models publish are GPQA Diamond, MMMLU and Global-MMLU, so I ran Jev on those. The older benchmarks are here too, but their comparison models are older.

The frontier reaches its GPQA numbers with extended thinking. Jev, with no reasoning at all, sits roughly where reasoning models of late 2024 sat. Chemistry was its weakest area at 66.7%, against 82.3% for physics and 72.4% for biology.

Published scores and Jev, by benchmark
GPT-6 Astra · thinking, any effort
96.0%
Gemini 3.1 Pro · thinking high
94.3%
Claude Fable 5.1 · max effort
93.7%
Kimi K3 · max
93.5%
Qwen3.8 Max · thinking
92.6%
DeepSeek V4.1 Flash · max effort
90.9%
Jev · one pass, no thinking
74.0%
Human PhD experts · GPQA paper, in their own field
about 65%
GPT-4 (2023) · GPQA paper, few-shot chain of thought
about 36%

198 graduate-level science questions with 4 options. Jev is averaged over 4 option orders. Bars start at 0%.

Accuracy and price on GPQA Diamond

Jev charges $0.042 per million input tokens and nothing for output. A GPQA question is about 535 tokens. The chart prices that same prompt at each model's list input price. Output and thinking tokens are left out, so the real gap is larger than shown.

GPQA Diamond accuracy against cost per question
70%80%90%100%$0.00001$0.0001$0.001$0.01cost per question (log scale)Jev 74.0%GPT-6 AstraGemini 3.8 FlashGrok 4.6Claude Fable 5.1Kimi K3DeepSeek V4 Pro 0813GLM-5.3DeepSeek V4.1 Flash

Jev: 74.0% on GPQA Diamond. Measured cost $0.000022 per question, median answer time 0.20 s.

One point per model released or updated from July to September 2026, using the lab's own number where there is one. Other models' costs are estimates. Tap a point or a name for details.

How often its confidence is right

Every answer comes with a probability. I took all 4,036 LSAT answers and grouped them by how confident Jev said it was, then checked how often each group was right. Of the answers where it said about 70%, about 74% were right. Of the ones where it said about 90%, about 92% were right.

That makes the number usable. If you only accept answers where it is at least 80% sure, it answers 75% of the questions and gets 97.3% of those right. At a 70% bar it answers 80% and gets 96.5% right. Most language models are overconfident, so their stated confidence cannot be used this way.

Stated confidence against actual accuracy, LSAT
0%0%25%25%50%50%75%75%100%100%perfectly honestSaid 27.0%, right 20.3% of 74 answersSaid 34.9%, right 39.0% of 187 answersSaid 44.4%, right 34.9% of 175 answersSaid 54.2%, right 54.7% of 181 answersSaid 64.5%, right 74.4% of 191 answersSaid 74.9%, right 85.0% of 200 answersSaid 85.0%, right 92.0% of 238 answersSaid 98.5%, right 97.7% of 2,790 answershow sure Jev said it washow often it was right

Jev answers 75.0% of the 4,036 LSAT answers and gets 97.3% of those right. The rest go to a slower model or a person.

Each dot is a group of answers with similar confidence. The big dot holds 2,790 answers above 90%. Expected calibration error is 0.027.

What it does when people disagree

ChaosNLI is a set of sentence pairs. The task is to say whether the second sentence follows from the first, contradicts it, or neither. Each pair was labeled by 100 people. On some pairs nearly everyone agrees. On others they split, because the pair really is ambiguous.

When people disagreed, Jev became less sure too. Its average top probability was 0.95 on pairs where people agreed, 0.88 on partly split pairs and 0.78 on badly split ones. A model trained to sound right tends to pick a side even when there is no right side. This is the best evidence I have that TypeSafe's training method, which they call RLCD, does what they say.

Human votes and Jev's probabilities for 20 of the 200 pairs
people split
partly split
people agreed

First sentence

Several people stopped by the side of the road are watching what is happening in the woods.

Second sentence

People are talking to eachother

follows
4 of 100
0%
neither
78 of 100
100%
contradicts
18 of 100
0%

Gray: votes from 100 annotators. Orange: Jev's probability for each label.

Within each agreement level the pairs span the range of Jev's confidence. Pair 8, shown first, is a pair where 78 people said neither, 18 contradiction and 4 follows. Jev put all its probability on neither, so it does not always hedge.

Following rules it has never seen

The LSAT and MMLU are public, so a high score may be memory. To test reasoning on something Jev cannot have seen, code generated problems with made-up rules and nonsense words, such as "a drishev is any red round object". Stated rules were essentially solved. That held for rules that contradict common sense, exceptions, rules revised later in the text, and one relevant rule hidden among 400.

Long chains are where it fails. Each category requires the one before it plus a named mark, and Jev has to find the deepest category that applies. Accuracy falls from 93% at 2 links to 8% at 16, close to the 6% chance rate. Its confidence stops tracking accuracy here: at 16 links it still reports 43%. It judges from the overall look of the input and does not walk the chain link by link.

Accuracy by chain length
0%25%50%75%100%2345681216links in the chain2 links: 93% right, said 83% sure, chance 33%3 links: 80% right, said 76% sure, chance 25%4 links: 75% right, said 70% sure, chance 20%5 links: 48% right, said 60% sure, chance 17%6 links: 27% right, said 50% sure, chance 14%8 links: 20% right, said 46% sure, chance 11%12 links: 12% right, said 44% sure, chance 8%16 links: 8% right, said 43% sure, chance 6%accuracy 93%stated confidence 43%accuracy 8%chance
60 generated problems per length. Dotted gray: Jev's average stated confidence. Dashed: chance.

Working out an unstated rule from labeled examples levels off in the 70s. Arithmetic before a comparison is close to guessing. Code can do both jobs: it can walk a chain one link per call, and it can add.

Accuracy by language

On MMMLU, Jev averaged 83.9% over 14 languages and 91.4% in English on the same questions. The loss from translation is small for European and East Asian languages. Most of it is in Swahili (75.4%) and Yoruba (60.8%). Global-MMLU-Lite shows the same pattern over 23 languages.

Jev's accuracy by language
English (control)
91.4%
Portuguese
90.2%
Italian
89.4%
Spanish
89.2%
French
88.2%
German
87.8%
Indonesian
87.2%
Chinese
86.8%
Korean
86.6%
Japanese
86.4%
Arabic
84.6%
Hindi
81.8%
Bengali
80.6%
Swahili
75.4%
Yoruba
60.8%
Thin line: English on the same questions, 91.4%

Bars start at 50%. Chance is 25%.

Chess

Search did not help in chess, because Jev's judgment of chess positions is poor to begin with. In one pass it picks Stockfish's best move about 23% of the time and blunders on about 45% of moves. Against a 1350-rated engine it won none of 10 games. How the board was written down made no measurable difference.

One move per position, 200 positions, judged by Stockfish at depth 12
Board written asBest moveBest in top 3Within 30 cpBlunders
8×8 ASCII grid21.5%38.0%29.5%49.5%
FEN string23.5%39.0%31.0%44.5%
Piece list in prose23.5%42.5%33.0%44.0%
Full games against Stockfish at 1350 Elo
How Jev playedWonDrawnLostJev calls per move
Top pick, no search0241
MCTS, 24 simulations01345

There is also a technical cause. Through the gateway, probabilities come back rounded to two decimals. With 30 or more legal moves, most moves read 0.01 to 0.03, so the ranking is nearly flat and search has little to work with.

Consistency checks

A model with honest probabilities should also give the same answer when you ask the same thing in a different way. Jev mostly does. The clear exception is negation. Asked whether a statement is true and, separately, whether it is false, its two probabilities should add to 1. They miss by 0.13 on average. TypeSafe lists negation as a known weakness.

CheckResult
Same question, 3 option orders × 3 phrasings: answer changed1.0%
Which of two texts is more formal, 200 triples: preference cycles (A > B > C > A)0.5%
"Is it true?" plus "is it false?": average distance from 10.13
Question asked alone vs. with another question: average shift0.02

I also added junk and planted answers to 200 LSAT questions. Sixteen unrelated paragraphs cost 2 points. A note saying "the correct answer is" with a wrong letter cost 3 points, and Jev followed the planted letter 6% of the time. A fake "SYSTEM: ignore the passage" line was followed 3% of the time. Its confidence dropped in those conditions.

200 LSAT questions under each condition
ConditionAccuracyMean confidence
Clean94.0%0.95
16 unrelated paragraphs added92.0%0.92
Planted wrong answer91.0%0.89
Fake system instruction94.0%0.91

What it knows

TypeSafe has said almost nothing about what is inside Jev. Its knowledge scores say there is a large pretrained language model. It can pick a missing word out of 255 candidates 71% of the time. Its knowledge of rare entities drops off the way a language model's does. It answered questions about events through the end of 2025 perfectly and missed some from 2026, so its training data ends around early 2026.

TestQuestionsJevChance
MMLU, 10 per subject57091.1%25%
MMLU-Pro14081.4%11%
PopQA, rare and common entities1,00069.7%25%
PopQA, obscure entities only62.5%25%
PopQA, famous entities only92.8%25%
Dated events, 2022 to 20264895.8%25%
Missing word, 255 candidates20071.0%0.4%

What it is good for

Jev is a fast judge with honest confidence. It is weak at anything that needs several dependent steps. The practical design puts the steps in code or search, and uses Jev's confidence to decide when to hand a case to a slower model. I tried that with a browser agent in a second post.

Methods and caveats

How Jev was run+

All calls went to typesafe-ai/jev through Vercel AI Gateway between 19 and 25 September 2026. Each question was one Choice call: the passage and question as the input, the answer options as the choices. No examples, no retries for correctness. The earlier study cost about $1 in Jev calls. GPQA cost 1.8 cents, MMMLU 15 cents and Global-MMLU-Lite 10 cents.

TypeSafe's own published accuracy numbers are scored against what GPT-6 Astra and Claude Fable 5.1 agree on. That measures agreement with those models. Every number here is scored against the dataset's real labels instead.

Published numbers and their settings+

No frontier model was run for this post. Every comparison is a published number from a lab page, model card, system card or an evaluator (Artificial Analysis, Vals, Epoch AI, llm-stats). Anthropic no longer publishes GPQA, so the Claude GPQA number is an Artificial Analysis run at max effort. Most comparison numbers come from models that generated reasoning tokens before answering. Jev did not.

Costs for other models in the price chart are estimates: 535 prompt tokens times each model's list input price. Output and thinking tokens are not counted.

Sample sizes and intervals+

GPQA Diamond: all 198 questions, each in 4 option orders; 95% interval 70.8% to 76.9%. Jev chose the same option across orderings on 75% of questions. MMLU is a 570-question subset and MMLU-Pro a 140-question subset, so a difference of a few points against a leaderboard row is within noise. The LSAT run is all 1,009 AGIEval questions in 4 option orders.

MMMLU: some translations reorder questions while reusing ids, so six affected subjects were dropped in every language and 500 paired questions were sampled per language. It is a filtered sample, not the official full set. Global-MMLU-Lite used 200 questions in each of 23 languages.

Contamination+

AGIEval LSAT and MMLU are old and public. Jev has probably seen them in training, so those scores show what it knows, not how it would do on a new exam. The invented-rule problems are generated from a seed and cannot have been seen.

ChaosNLI metric+

Jev's Jensen-Shannon distance to the human vote distribution is 0.21 on 200 pairs. Fine-tuned 2020 models scored 0.22 to 0.24, 2026 reasoning models about 0.12 after chain of thought, and a second set of human annotators 0.06. An earlier draft of my notes reported 0.06 for Jev, which was the divergence, not the distance.

Search and chess details+

Game of 24: the same 30 seeded puzzles in every row. Beam search used width 5 and 157 calls per puzzle. The random-judge row comes from the run notes, not a separate result file. Chess games were 6 with no search and 4 with MCTS, too few to measure a difference.