Benchmark · October 2026
Cheaper on every task, by 1.1× to 3.6×.
One AI agent ran eight real tasks on live sites, three times through Naoba and three times through the browser built into the Claude desktop app. Same model, same accounts, the two runs of a pair at the same moment. Each task is compared on its own: how much a set of them costs depends on which sites you use.
Cost
1.1×–3.6× cheaper on every task
The more steps a task takes, the more Naoba saves. Median of three runs per task, at the model's list price.
Accuracy
7 of 8 tasks right as often or more
Better on 3, the same on 4, worse on 1 (YouTube). 1 of Naoba's misses and 2 of the built-in browser's came from a gap in the checker on Amazon.
Speed
5 of 8 tasks faster
Each step takes longer with Naoba, because it carries a whole scenario; the fewer steps win on most tasks. The same on Gmail, slower on Google and YouTube.
Steps
1.1×–3.6× fewer on every task
A step is one answer from the model followed by one tool call. Fewer steps is where the saving comes from.
Cost per task
The longer the task, the wider the gap.
Median cost of three runs at the model's list price. A short read costs about the same either way. On a long task the built-in browser takes one action per call and the agent pays to reread the whole conversation each time; with Naoba it writes several steps into one call.
NaobaBuilt-in browser
Steps per task
Fewer steps on every task.
A step is one answer from the model followed by one tool call. Fewer steps is where the saving comes from.
NaobaBuilt-in browser
Time per task
Faster on 5 tasks, slower on 2.
With Naoba the agent writes a whole scenario before each call, so each step takes longer. Where a task needs many steps, the fewer steps win; on Google and YouTube the built-in browser was faster.
NaobaBuilt-in browser
Every run
48 runs, 8 of them wrong.
A run is right when every checked fact in the agent's report matches the live site, read again by the checker right after the run. ✓ is right, ✕ is wrong; hover a mark for its score.
Triage a mailbox
Gmail, signed in, in Russian: list what matters, with a conversation that hides a second alert, a subject that contradicts its body and a relative deadline.
Find a book and its real cost
Amazon: pick the right edition among look-alikes, and read price, delivery and import charges from a popover.
Find somewhere to stay
Booking.com: a city, dates, filters and a price sort, then the three cheapest places with their real totals.
Look three things up
Google: a cookie wall, links that hide where they go, a release candidate next to the release, and a currency conversion.
Report on a project
GitHub: the current release among pre-releases, the most argued bugs, how one fix landed, and recent bug counts.
Check a talk before watching
YouTube: the real recording, its length, what it covers, where a part starts, and the channel's newest video.
Digest a discussion
Hacker News: the most argued stories, where in a thread the argument is, and what the site said over the years.
Read a long article
Wikipedia: facts from an infobox, a table, a paragraph, the page history and another language edition.
What went wrong
- Find a book and its real cost, Naoba, run 3 (score 78). After the run the checker could not read Amazon's own price breakdown, so it could not confirm the price the agent reported and counted it as wrong. This is a gap in the checker, not necessarily in the answer.
- Find a book and its real cost, Built-in browser, run 2 (score 89). After the run the checker could not read Amazon's own price breakdown, so it could not confirm the price the agent reported and counted it as wrong. This is a gap in the checker, not necessarily in the answer.
- Find a book and its real cost, Built-in browser, run 3 (score 78). After the run the checker could not read Amazon's own price breakdown, so it could not confirm the price the agent reported and counted it as wrong. This is a gap in the checker, not necessarily in the answer.
- Find somewhere to stay, Naoba, run 3 (score 38). Took a wider destination than the city from the search box Booking had pre-filled, and listed 524 places instead of 437. Both browsers gave this same answer once.
- Find somewhere to stay, Built-in browser, run 2 (score 38). Took a wider destination than the city from the search box Booking had pre-filled, and listed 528 places instead of 437. Both browsers gave this same answer once.
- Find somewhere to stay, Built-in browser, run 3 (score 75). Ranked by the price a signed-out guest pays and left a members-only price out of the three cheapest.
- Look three things up, Built-in browser, run 1 (score 80). Left the currency conversion out of the report.
- Check a talk before watching, Naoba, run 2 (score 90). Read "11 days ago" and gave the newest video's date one day late.
1 of Naoba's misses and 2 of the built-in browser's are that Amazon checker gap. The results stand as they were scored: nothing was rescored or dropped after the runs.
How it was run
One agent, two browsers, nothing else different.
The same agent
Claude Sonnet 5 at medium effort, as a subagent of a Claude desktop session, the only way both browsers can be reached on equal terms. No instruction files and no memory loaded.
The same task
The two prompts differ only in the paragraph that names the browser's tools. The agent gets the task, not a hint about how to do it.
The same moment
The two runs of a pair started together on two sessions, which swapped browsers every round. All three rounds of a task fell on one day.
The same scoring
One checker for both, reading the moving facts live after each run. Checks only one browser could be asked, such as declining a cookie question, are left out of both scores.
Limitations
- One Mac, one network, one Google account, one model. A different model may change the balance.
- Live sites move between runs. The checker rereads what can move, but a site can still answer two runs differently.
- Both browsers were warmed once before the runs, and each cookie question was declined then. Neither started cold.
- Prices are the model's list prices. The Claude desktop app writes only the first few output tokens of most replies into its record, so the output part of the cost is estimated from the length of each reply and of its hidden reasoning, with a formula checked on the 72 replies whose true count survives (typical error 3%, worst 15%). Input and cache tokens, most of the bill, are exact. The data file carries both the estimated and the recorded cost of every run.
- Naoba's tools reach the agent on demand while the built-in browser's are loaded in full, as each comes by default.
- The desktop app updated its Claude Code during the cycle: the last two tasks ran on a newer version, both browsers alike.
Every run, its checks and its cost: benchmark-2026-10.json.
The numbers as a table
| Task | Browser | Right | Cost | Time | Steps | Browser calls |
|---|---|---|---|---|---|---|
| Triage a mailbox | Naoba | 3 of 3 | $1.04 | 6.5 min | 33 | 29 |
| Triage a mailbox | Built-in browser | 3 of 3 | $2.94 | 6.8 min | 120 | 117 |
| Find a book and its real cost | Naoba | 2 of 3 | $0.93 | 4.0 min | 25 | 21 |
| Find a book and its real cost | Built-in browser | 1 of 3 | $2.06 | 5.5 min | 66 | 69 |
| Find somewhere to stay | Naoba | 2 of 3 | $0.52 | 3.5 min | 23 | 20 |
| Find somewhere to stay | Built-in browser | 1 of 3 | $1.89 | 4.3 min | 72 | 70 |
| Look three things up | Naoba | 3 of 3 | $0.89 | 3.7 min | 36 | 32 |
| Look three things up | Built-in browser | 2 of 3 | $1.09 | 3.0 min | 52 | 49 |
| Report on a project | Naoba | 3 of 3 | $0.54 | 2.7 min | 22 | 18 |
| Report on a project | Built-in browser | 3 of 3 | $0.83 | 3.5 min | 38 | 35 |
| Check a talk before watching | Naoba | 2 of 3 | $0.54 | 4.6 min | 26 | 22 |
| Check a talk before watching | Built-in browser | 3 of 3 | $0.59 | 2.1 min | 29 | 26 |
| Digest a discussion | Naoba | 3 of 3 | $0.40 | 2.0 min | 17 | 14 |
| Digest a discussion | Built-in browser | 3 of 3 | $0.48 | 2.2 min | 22 | 21 |
| Read a long article | Naoba | 3 of 3 | $0.39 | 1.8 min | 16 | 12 |
| Read a long article | Built-in browser | 3 of 3 | $0.48 | 2.0 min | 21 | 17 |
Medians of three runs.