Skip to content
← Blog
AIAgentsEvals

A browser agent built on Jev

Jev is a model that cannot write text. It picks one option from a list and says how sure it is, in about a quarter of a second. I tested the model first. This post is about building a browser agent around it.

The harness reads the page and lists what can be clicked, typed into or scrolled. Jev picks the element and the action at every step, and also decides when to stop. A small language model (gpt-5.4-mini) writes the text that gets typed and the one-line final answer, because Jev cannot. I ran it on Online-Mind2Web, which is 300 tasks on 136 live websites.

Scorecard
Online-Mind2Web success34.7%Published: Browser Use Cloud 97%, own Claude-based judge. Claude 4.5: 55% by human judges.Note: 35 of the 101 tasks judged so far. 198 of 299 have no verdict yet. Counting those as failures gives 11.7%.
Time per task27.3 sPublished: No leaderboard entry publishes time per task.Note: Median, headed Chrome. About 1.2 s per step including page loads.
Cost per task$0.0053Published: Not published.Note: Jev plus the small model that types. Judging is extra.
Offline Mind2Web step success50.7%Published: MindAct, fine-tuned: 52.0%. GPT-4 with 3 examples and a ranker: 36.2%.Note: Held-out steps, up from 7.3% before tuning. Our metric does not check typed text; theirs does.

The agent is far behind the frontier agents on success, roughly where the weaker agents in the 2025 paper sat. It finishes a task in about half a minute for about half a cent. Most of its failures come from choosing the wrong element and not knowing when to stop.

Where it started: recorded web pages

Before going live I tuned the harness on Mind2Web, a set of recorded web tasks split into single steps. For each step the harness lists page elements and Jev chooses one. A coding agent (Codex running GPT-6 Astra) was allowed to edit one harness file, and a fixed evaluator scored each edit on 200 tuning steps. A separate set of 300 steps from other tasks was held out and run by hand after each round.

The held-out score went from 7.3% to 34.0%, 37.3%, then 50.7%. Showing Jev every element in batches, and choosing click, type or select by a rule, was worth about 27 points. The next biggest gain came from a bug in my data: the loader had dropped the visible text of page elements. Fixing it was worth about 13 points. Everything else added about 3.

Offline Mind2Web step success across three rounds of tuning
0%10%20%30%40%50%60%element text fixeddev run 0: 7.5%. baseline: 2-stage hierarchical, groups of 25, op asked with elementdev run 1: 8.5%. 01 retry intermittent Unknown model failures; retain baseline selectiondev run 2: 10.5%. 02 render semantic IDs classes href and current input value to expose missing text cluesdev run 3: 11.5%. 03 recognize dataset aria_label alt text_value attribute spellingsdev run 4: 23.0%. 04 full-coverage batched tournament in groups of 100 avoids blind group summariesdev run 5: 23.0%. 05 bound tournament prompt batches and compact rendering to fix large-pool failuresdev run 6: 23.0%. 06 retry transient service failures in bounded tournamentdev run 7: 29.5%. 07 derive operation from chosen tag and input type with date-picker exceptionsdev run 8: 29.5%. 08 prioritize accessible labels before verbose CSS classes in compact descriptionsdev run 9: 28.0%. 09 cheap lexical top-80 single Choice tests accuracy-cost tradeoffdev run 10: 29.0%. 10 widen cheap lexical shortlist to 160 to recover excluded targetsdev run 11: 32.0%. 11 spatially ordered tournament groups compare nearby page elementsdev run 12: 30.0%. 12 widen single-call lexical shortlist to 240 near Choice capacitydev run 13: 32.0%. 13 carry top 3 probabilities from each spatial batch to recover first-round lossesdev run 14: 29.5%. 13 actual top 3 carryover; prior 13-labelled run was unchanged baseline after command failuredev run 15: 32.0%. 14 carry top 5 per batch to increase survival despite final-choice dilutiondev run 16: 33.5%. 15 top 5 carryover with fuller labelled tag role and attributes in final Choicedev run 17: 31.0%. 16 fuller tag role and labelled attributes in first round as well as finaldev run 18: 26.0%. 17 cheap top 80 with full-attribute overlap and broader interactive priordev run 19: 32.0%. 18 cheap top 240 with broader interactive prior tests high shortlist recalldev run 20: 32.0%. 19 fan-out Noul on top 8 final candidates combined with square-root Choice probabilitydev run 21: 35.0%. 20 enrich interactive candidates with contained visual attribute clues as an ancestry proxydev run 22: 34.0%. 21 collapse identical compact descriptions and choose an interactive representativedev run 23: 29.5%. 22 task-progress Choice conditions deduplicated tournament on a code-generated next sub-goaldev run 24: 31.5%. 23 last-action priors favor follow-up clickable targets and avoid repeating a typed fielddev run 25: 35.5%. 24 confidence gate 0.5 uses cheap top 240 then full geometry tournament only on uncertaintydev run 26: 35.0%. 25 confidence gate 0.5 falls back to deduplicated geometry tournament to cut costdev run 27: 33.0%. 26 one-call top 240 reuses contained visual context and fuller finalist descriptionsdev run 28: 31.5%. 27 state wording ablation keeps only last four completed actions in confidence-gated agentdev run 29: 35.0%. 28 instruction ablation emphasizes immediate interactive target and avoiding completed actionsdev run 30: 9.0%. WITH ELEMENT TEXT: agent-00-0.075-baseline unchangeddev run 31: 38.5%. WITH ELEMENT TEXT: agent-09-0.2800 unchangeddev run 32: 48.0%. WITH ELEMENT TEXT: agent-25-0.3500 unchangeddev run 33: 39.0%. 29 R3 text context labels lexical top80 interactive prior compact renderingdev run 34: 32.0%. 30 R3 top40 tests aggressive lexical pruningdev run 35: 42.5%. 31 R3 top160 recovers lexical omissions within a compact single calldev run 36: 39.0%. 32 R3 shorter text and rendering fit 100 candidates in single-call budgetdev run 37: 45.0%. 33 R3 top160 then five richer finalists with ancestor context and pathdev run 38: 44.5%. 34 R3 eight richer finalists tests recovery from rounded probability tiesdev run 39: 44.0%. 35 R3 confidence 0.5 skips rich final when compact choice is confidentdev run 40: 45.0%. 36 R3 confidence 0.7 trades more rich finals for fewer premature exitsdev run 41: 48.0%. 37 R3 DOM order and ancestor path regions cover full pool before five rich finalistsdev run 42: 50.5%. 38 R3 collapse nearby matching parent-child descriptions preferring interactive ancestordev run 43: 51.0%. 39 R3 parent-child collapse instead prefers labelled interactive controlsdev run 44: 45.0%. 40 R3 parallel operation choice distinguishes input focus from typing with native-select guarddev run 45: 41.0%. 41 R3 DOM-region tournament on lexical top160 with shorter text and two survivors per groupdev run 46: 50.0%. 42 R3 compress full-pool DOM tournament descriptions while keeping rich finaldev run 47: 38.0%. 43 R3 brief ancestor context supplies missing labels for textless controls in cheap pathdev run 48: 48.0%. 44 R3 richer descriptions before narrowing DOM survivors to final fivedev run 49: 45.0%. 45 R3 final instruction distinguishes immediate action and candidate text from ancestor contextdev run 50: 39.0%. 46 R3 slightly wider single call with tighter total descriptions targets 3000 token ceilingdev run 51: 38.5%. 47 R3 labelled parent-child collapse before cheap top80 lexical rankingdev run 52: 44.5%. 48 R3 labelled parent-child collapse before top160 and confidence-gated five-way finaldev run 53: 50.0%. R3 final verification: selected agent42 with inactive experiment branches removed7.3%34.0%37.3%50.7%tuning runs, in order
one tuning run on 200 dev steps best held-out score so far, 300 steps
The flat stretches are rounds where most ideas did not help. After the text fix, the round-two harness moved from 35% to 48% on the tuning steps without any code change.

Headless browser or real Chrome

I ran all 300 live tasks twice. A headless browser is fingerprinted and refused by Akamai and Cloudflare bot walls on 26 sites, so it never gets a turn there. Real Chrome on macOS passes most of them. Headed Chrome is also what the published agents use, usually through remote browser services.

The same 300 tasks, two browsers
headless browser headed real Chrome
Blocked by a bot wall
20.4% (61/299)
4.0% (12/299)
Ended with an answer (complete, or answered on exit)
12.4% (37/299)
21.4% (64/299)
Judged success, among tasks the judge scored
23.4% (18/77)
34.7% (35/101)
HeadlessHeaded Chrome
Median time per task13.0 s27.3 s
Median time per step1.19 s1.17 s
Agent cost per task$0.0037$0.0053

The headed run takes about twice as long per task. Pages actually load there, so the agent takes more steps. The time per step is the same.

Success by difficulty

Online-Mind2Web labels each task easy, medium or hard. Among judged tasks, Jev's agent succeeded on 19 of 36 easy tasks, 13 of 41 medium and 3 of 24 hard. The dark part of each bar is tasks with no verdict yet.

Headed run, by task difficulty
judged success judged failure no verdict yet
Easy19/36 scored (53%) · 80 tasks
Medium13/41 scored (32%) · 140 tasks
Hard3/24 scored (13%) · 79 tasks

How tasks ended

A run stops when the agent says the task is done, says it is impossible, or hits the 25-step limit. When it stops without having said done, the small model still writes an answer from what it saw. That exit answer accounts for 8 of the 35 successes.

The two big groups are failures to finish. 115 tasks ended with the agent calling the task impossible and 95 hit the step limit. Most steps in those runs were scrolling or repeated clicks that did not move the task forward.

Headed run, by how the task ended
judged success judged failure no verdict yet
Agent said the task was complete12/19 scored (63%) · 37 tasks
Answered when it stopped8/14 scored (57%) · 27 tasks
Hit the 25-step limit10/32 scored (31%) · 95 tasks
Agent said the task was impossible4/31 scored (13%) · 115 tasks
Blocked by a bot wall0/1 scored (0%) · 12 tasks
Error1/4 scored (25%) · 11 tasks
Hit the time limit0/0 scored · 2 tasks

Why a live benchmark needs a judge

On recorded pages there is an answer key: the element a person clicked. On the live web there is none, because pages change and a task can often be finished in more than one way. Someone has to look at what the agent did and decide.

Online-Mind2Web uses WebJudge. A language model reads the task, pulls out the key points, rates the agent's screenshots, and reads its action list, then says success or failure. I used gpt-5.4-mini as the judge, which is cheaper and weaker than the Claude-class judges most leaderboard entries use. Judges disagree with each other and with people by 10 to 15 points on the same agent, in both directions.

Same agent, two judges
human judges WebJudge
Yutori Navigator
human 78.7% · WebJudge 64.7% · gap 14.0 points
Operator (paper)
human 61.3% · WebJudge 71.8% · gap 10.5 points
Claude 4.5
human 55.0% · WebJudge 59.3% · gap 4.3 points
40%50%60%70%80%90%
Published results that report both human and WebJudge scores.

My judging is also unfinished. Gateway credits ran out partway, so 101 of the 299 headed tasks have a verdict and 198 do not. On the judged tasks the agent succeeded 35 times, which is 34.7%. If every unjudged task were a failure, the rate would be 11.7%. The judged tasks are the first 101 the judge reached, not a chosen sample, and 101 tasks gives about ±9 points.

Against the published leaderboard

Scores from different judges are not directly comparable, so the judge is listed with each entry. Even allowing for that, the gap to the top is large.

Online-Mind2Web success rates
Browser Use Cloud (bu-max) Mar 202697.0%
Judge: custom Claude-based judge
GPT-5.4 native computer use Mar 202693.0%
Judge: screenshot-based judge
ABP + Claude Opus 4.6 Mar 202690.5%
Judge: per-task results published
UI-TARS-2 Sep 202588.2%
Judge: standard Online-Mind2Web
Yutori Navigator Nov 202578.7%
Judge: human 78.7%, WebJudge 64.7%
Operator (paper) Apr 202561.3%
Judge: human 61.3%, WebJudge 71.8%
Claude 4.5 Nov 202555.0%
Judge: human 55%, WebJudge 59.3%
Jev harness, headed (this run) Sep 202634.7%
Judge: WebJudge-style, gpt-5.4-mini; 101 of 299 tasks scored. White tick: 11.7% if every unjudged task counts as a failure.
Published entries from the Online-Mind2Web leaderboard and paper. Jev's bar uses the 101 judged tasks.

Time per task

The median task took 27.3 seconds and the mean 48 seconds. A few tasks ran much longer, including one that hung on remax.com in both runs. An average step spent about 1.1 seconds on Jev calls through the gateway and about 0.8 seconds in the browser.

Seconds per task, headed run
040800 to 10 s: 62 tasks62010 to 20 s: 52 tasks521020 to 30 s: 45 tasks452030 to 45 s: 57 tasks573045 to 60 s: 33 tasks334560 to 90 s: 25 tasks256090 to 120 s: 7 tasks790120 to 240 s: 16 tasks16120240 or more s: 2 tasks2240seconds per task (uneven bins)
299 tasks. Nobody on the leaderboard publishes time per task. The only published speed figure is Navigator's own claim of 3.3 times faster per step than Claude 4.5.

What limits it now

Element choice and stopping are the main problems. The typed text still comes from a small language model, with Jev picking the field. Bot walls block 4% of tasks even in real Chrome, on 6 sites. And the judge often returned "unsure" when it could not verify a claim from screenshots.

The 198 unjudged tasks still need a verdict before the success rate is final.

Methods and caveats

The live run+

Run 2026-09-25 on 300 Online-Mind2Web tasks (80 easy, 141 medium, 79 hard), seed 1, 4 tasks at a time, at most 25 steps and 240 seconds per task. Jev (typesafe-ai/jev through Vercel AI Gateway) chose every element and every control action: act, scroll, back, complete, impossible. gpt-5.4-mini wrote typed values and final answers. One task (remax.com) hung in both runs, so each run finished 299 tasks.

Agent cost was about $1.60 for 299 headed tasks and judging about $2.20 for 101 tasks. The headless run and its partial judging cost about $5.30.

Judging+

WebJudge-style evaluation with gpt-5.4-mini: extract key points from the task, rate each screenshot for relevance, then judge the final state and action history. These automated verdicts are not directly comparable to human judgments. Headed: 35/101 judged success, 198 not judged. Headless: 18/77 judged success, 222 not judged. The headless judging stopped even earlier, so its rate rests on fewer tasks.

The offline Mind2Web tuning+

The evaluator scored each harness version on Mind2Web steps using every candidate element on the page (median 419 per step), capped Jev at 12 calls per step and the whole loop at $8. The 177 tasks were split in half by task. Tuning used 200 steps from one half; the other half was held out, blocked from the coding agent, and run three times in total. The winning file was checked for step ids, site names and file reads.

Step success here counts the right element and operation. The Mind2Web paper also requires the typed text to be right, so its numbers are stricter. 300 steps gives an interval of roughly ±6 points. The loader fix changed what Jev could see, so scores before and after it are not measured on the same input. A cheaper variant scores 49.0% on held-out steps with 40% of the tokens and a sixth of the time.

Leaderboard sources+

Published entries come from the Online-Mind2Web leaderboard as aggregated by steel.dev (updated 2026-06-29) and the Online-Mind2Web paper. Each entry is listed with the judge its authors used. No other agent was rerun for this post.