Naoba

Benchmark · October 2026

Cheaper on every task, by 1.1× to 3.6×.

One AI agent ran eight real tasks on live sites, three times through Naoba and three times through the browser built into the Claude desktop app. Same model, same accounts, the two runs of a pair at the same moment. Each task is compared on its own: how much a set of them costs depends on which sites you use.

Cost

1.1×–3.6× cheaper on every task

The more steps a task takes, the more Naoba saves. Median of three runs per task, at the model's list price.

Accuracy

7 of 8 tasks right as often or more

Better on 3, the same on 4, worse on 1 (YouTube). 1 of Naoba's misses and 2 of the built-in browser's came from a gap in the checker on Amazon.

Speed

5 of 8 tasks faster

Each step takes longer with Naoba, because it carries a whole scenario; the fewer steps win on most tasks. The same on Gmail, slower on Google and YouTube.

Steps

1.1×–3.6× fewer on every task

A step is one answer from the model followed by one tool call. Fewer steps is where the saving comes from.

Cost per task

The longer the task, the wider the gap.

Median cost of three runs at the model's list price. A short read costs about the same either way. On a long task the built-in browser takes one action per call and the agent pays to reread the whole conversation each time; with Naoba it writes several steps into one call.

NaobaBuilt-in browser

Steps per task

Fewer steps on every task.

A step is one answer from the model followed by one tool call. Fewer steps is where the saving comes from.

NaobaBuilt-in browser

Time per task

Faster on 5 tasks, slower on 2.

With Naoba the agent writes a whole scenario before each call, so each step takes longer. Where a task needs many steps, the fewer steps win; on Google and YouTube the built-in browser was faster.

NaobaBuilt-in browser

Every run

48 runs, 8 of them wrong.

A run is right when every checked fact in the agent's report matches the live site, read again by the checker right after the run. ✓ is right, ✕ is wrong; hover a mark for its score.

Triage a mailbox

Gmail, signed in, in Russian: list what matters, with a conversation that hides a second alert, a subject that contradicts its body and a relative deadline.

Naoba✓✓✓
Built-in✓✓✓

Find a book and its real cost

Amazon: pick the right edition among look-alikes, and read price, delivery and import charges from a popover.

Naoba✓✓✕
Built-in✓✕✕

Find somewhere to stay

Booking.com: a city, dates, filters and a price sort, then the three cheapest places with their real totals.

Naoba✓✓✕
Built-in✓✕✕

Look three things up

Google: a cookie wall, links that hide where they go, a release candidate next to the release, and a currency conversion.

Naoba✓✓✓
Built-in✕✓✓

Report on a project

GitHub: the current release among pre-releases, the most argued bugs, how one fix landed, and recent bug counts.

Naoba✓✓✓
Built-in✓✓✓

Check a talk before watching

YouTube: the real recording, its length, what it covers, where a part starts, and the channel's newest video.

Naoba✓✕✓
Built-in✓✓✓

Digest a discussion

Hacker News: the most argued stories, where in a thread the argument is, and what the site said over the years.

Naoba✓✓✓
Built-in✓✓✓

Read a long article

Wikipedia: facts from an infobox, a table, a paragraph, the page history and another language edition.

Naoba✓✓✓
Built-in✓✓✓

What went wrong

  • Find a book and its real cost, Naoba, run 3 (score 78). After the run the checker could not read Amazon's own price breakdown, so it could not confirm the price the agent reported and counted it as wrong. This is a gap in the checker, not necessarily in the answer.
  • Find a book and its real cost, Built-in browser, run 2 (score 89). After the run the checker could not read Amazon's own price breakdown, so it could not confirm the price the agent reported and counted it as wrong. This is a gap in the checker, not necessarily in the answer.
  • Find a book and its real cost, Built-in browser, run 3 (score 78). After the run the checker could not read Amazon's own price breakdown, so it could not confirm the price the agent reported and counted it as wrong. This is a gap in the checker, not necessarily in the answer.
  • Find somewhere to stay, Naoba, run 3 (score 38). Took a wider destination than the city from the search box Booking had pre-filled, and listed 524 places instead of 437. Both browsers gave this same answer once.
  • Find somewhere to stay, Built-in browser, run 2 (score 38). Took a wider destination than the city from the search box Booking had pre-filled, and listed 528 places instead of 437. Both browsers gave this same answer once.
  • Find somewhere to stay, Built-in browser, run 3 (score 75). Ranked by the price a signed-out guest pays and left a members-only price out of the three cheapest.
  • Look three things up, Built-in browser, run 1 (score 80). Left the currency conversion out of the report.
  • Check a talk before watching, Naoba, run 2 (score 90). Read "11 days ago" and gave the newest video's date one day late.

1 of Naoba's misses and 2 of the built-in browser's are that Amazon checker gap. The results stand as they were scored: nothing was rescored or dropped after the runs.

How it was run

One agent, two browsers, nothing else different.

The same agent

Claude Sonnet 5 at medium effort, as a subagent of a Claude desktop session, the only way both browsers can be reached on equal terms. No instruction files and no memory loaded.

The same task

The two prompts differ only in the paragraph that names the browser's tools. The agent gets the task, not a hint about how to do it.

The same moment

The two runs of a pair started together on two sessions, which swapped browsers every round. All three rounds of a task fell on one day.

The same scoring

One checker for both, reading the moving facts live after each run. Checks only one browser could be asked, such as declining a cookie question, are left out of both scores.

Limitations

  • One Mac, one network, one Google account, one model. A different model may change the balance.
  • Live sites move between runs. The checker rereads what can move, but a site can still answer two runs differently.
  • Both browsers were warmed once before the runs, and each cookie question was declined then. Neither started cold.
  • Prices are the model's list prices. The Claude desktop app writes only the first few output tokens of most replies into its record, so the output part of the cost is estimated from the length of each reply and of its hidden reasoning, with a formula checked on the 72 replies whose true count survives (typical error 3%, worst 15%). Input and cache tokens, most of the bill, are exact. The data file carries both the estimated and the recorded cost of every run.
  • Naoba's tools reach the agent on demand while the built-in browser's are loaded in full, as each comes by default.
  • The desktop app updated its Claude Code during the cycle: the last two tasks ran on a newer version, both browsers alike.

Every run, its checks and its cost: benchmark-2026-10.json.

The numbers as a table
TaskBrowserRightCostTimeStepsBrowser calls
Triage a mailboxNaoba3 of 3$1.046.5 min3329
Triage a mailboxBuilt-in browser3 of 3$2.946.8 min120117
Find a book and its real costNaoba2 of 3$0.934.0 min2521
Find a book and its real costBuilt-in browser1 of 3$2.065.5 min6669
Find somewhere to stayNaoba2 of 3$0.523.5 min2320
Find somewhere to stayBuilt-in browser1 of 3$1.894.3 min7270
Look three things upNaoba3 of 3$0.893.7 min3632
Look three things upBuilt-in browser2 of 3$1.093.0 min5249
Report on a projectNaoba3 of 3$0.542.7 min2218
Report on a projectBuilt-in browser3 of 3$0.833.5 min3835
Check a talk before watchingNaoba2 of 3$0.544.6 min2622
Check a talk before watchingBuilt-in browser3 of 3$0.592.1 min2926
Digest a discussionNaoba3 of 3$0.402.0 min1714
Digest a discussionBuilt-in browser3 of 3$0.482.2 min2221
Read a long articleNaoba3 of 3$0.391.8 min1612
Read a long articleBuilt-in browser3 of 3$0.482.0 min2117

Medians of three runs.