Jev is a model that cannot write text. It picks one option from a list and says how sure it is, in about a quarter of a second. I tested the model first. This post is about building a browser agent around it.
The harness reads the page and lists what can be clicked, typed into or scrolled. Jev picks the element and the action at every step, and also decides when to stop. A small language model (gpt-5.4-mini) writes the text that gets typed and the one-line final answer, because Jev cannot. I ran it on Online-Mind2Web, which is 300 tasks on 136 live websites.
The agent is far behind the frontier agents on success, roughly where the weaker agents in the 2025 paper sat. It finishes a task in about half a minute for about half a cent. Most of its failures come from choosing the wrong element and not knowing when to stop.
Where it started: recorded web pages
Before going live I tuned the harness on Mind2Web, a set of recorded web tasks split into single steps. For each step the harness lists page elements and Jev chooses one. A coding agent (Codex running GPT-6 Astra) was allowed to edit one harness file, and a fixed evaluator scored each edit on 200 tuning steps. A separate set of 300 steps from other tasks was held out and run by hand after each round.
The held-out score went from 7.3% to 34.0%, 37.3%, then 50.7%. Showing Jev every element in batches, and choosing click, type or select by a rule, was worth about 27 points. The next biggest gain came from a bug in my data: the loader had dropped the visible text of page elements. Fixing it was worth about 13 points. Everything else added about 3.
Headless browser or real Chrome
I ran all 300 live tasks twice. A headless browser is fingerprinted and refused by Akamai and Cloudflare bot walls on 26 sites, so it never gets a turn there. Real Chrome on macOS passes most of them. Headed Chrome is also what the published agents use, usually through remote browser services.
| Headless | Headed Chrome | |
|---|---|---|
| Median time per task | 13.0 s | 27.3 s |
| Median time per step | 1.19 s | 1.17 s |
| Agent cost per task | $0.0037 | $0.0053 |
The headed run takes about twice as long per task. Pages actually load there, so the agent takes more steps. The time per step is the same.
Success by difficulty
Online-Mind2Web labels each task easy, medium or hard. Among judged tasks, Jev's agent succeeded on 19 of 36 easy tasks, 13 of 41 medium and 3 of 24 hard. The dark part of each bar is tasks with no verdict yet.
How tasks ended
A run stops when the agent says the task is done, says it is impossible, or hits the 25-step limit. When it stops without having said done, the small model still writes an answer from what it saw. That exit answer accounts for 8 of the 35 successes.
The two big groups are failures to finish. 115 tasks ended with the agent calling the task impossible and 95 hit the step limit. Most steps in those runs were scrolling or repeated clicks that did not move the task forward.
Why a live benchmark needs a judge
On recorded pages there is an answer key: the element a person clicked. On the live web there is none, because pages change and a task can often be finished in more than one way. Someone has to look at what the agent did and decide.
Online-Mind2Web uses WebJudge. A language model reads the task, pulls out the key points, rates the agent's screenshots, and reads its action list, then says success or failure. I used gpt-5.4-mini as the judge, which is cheaper and weaker than the Claude-class judges most leaderboard entries use. Judges disagree with each other and with people by 10 to 15 points on the same agent, in both directions.
My judging is also unfinished. Gateway credits ran out partway, so 101 of the 299 headed tasks have a verdict and 198 do not. On the judged tasks the agent succeeded 35 times, which is 34.7%. If every unjudged task were a failure, the rate would be 11.7%. The judged tasks are the first 101 the judge reached, not a chosen sample, and 101 tasks gives about ±9 points.
Against the published leaderboard
Scores from different judges are not directly comparable, so the judge is listed with each entry. Even allowing for that, the gap to the top is large.
Time per task
The median task took 27.3 seconds and the mean 48 seconds. A few tasks ran much longer, including one that hung on remax.com in both runs. An average step spent about 1.1 seconds on Jev calls through the gateway and about 0.8 seconds in the browser.
What limits it now
Element choice and stopping are the main problems. The typed text still comes from a small language model, with Jev picking the field. Bot walls block 4% of tasks even in real Chrome, on 6 sites. And the judge often returned "unsure" when it could not verify a claim from screenshots.
The 198 unjudged tasks still need a verdict before the success rate is final.
Methods and caveats
The live run+
Run 2026-09-25 on 300 Online-Mind2Web tasks (80 easy, 141 medium, 79 hard), seed 1, 4 tasks at a time, at most 25 steps and 240 seconds per task. Jev (typesafe-ai/jev through Vercel AI Gateway) chose every element and every control action: act, scroll, back, complete, impossible. gpt-5.4-mini wrote typed values and final answers. One task (remax.com) hung in both runs, so each run finished 299 tasks.
Agent cost was about $1.60 for 299 headed tasks and judging about $2.20 for 101 tasks. The headless run and its partial judging cost about $5.30.
Judging+
WebJudge-style evaluation with gpt-5.4-mini: extract key points from the task, rate each screenshot for relevance, then judge the final state and action history. These automated verdicts are not directly comparable to human judgments. Headed: 35/101 judged success, 198 not judged. Headless: 18/77 judged success, 222 not judged. The headless judging stopped even earlier, so its rate rests on fewer tasks.
The offline Mind2Web tuning+
The evaluator scored each harness version on Mind2Web steps using every candidate element on the page (median 419 per step), capped Jev at 12 calls per step and the whole loop at $8. The 177 tasks were split in half by task. Tuning used 200 steps from one half; the other half was held out, blocked from the coding agent, and run three times in total. The winning file was checked for step ids, site names and file reads.
Step success here counts the right element and operation. The Mind2Web paper also requires the typed text to be right, so its numbers are stricter. 300 steps gives an interval of roughly ±6 points. The loader fix changed what Jev could see, so scores before and after it are not measured on the same input. A cheaper variant scores 49.0% on held-out steps with 40% of the tokens and a sixth of the time.
Leaderboard sources+
Published entries come from the Online-Mind2Web leaderboard as aggregated by steel.dev (updated 2026-06-29) and the Online-Mind2Web paper. Each entry is listed with the judge its authors used. No other agent was rerun for this post.