divyanshshukla.com

What 2,800 browser-agent runs taught me about planning

Divyansh Shukla

3 October 2026

Abstract

Plan-and-execute web agents matched a budget-matched ReAct baseline on 200 tasks with 83.6% fewer input tokens. What held up, what didn’t, and what to copy.

Most LLM web agents follow ReAct: the model picks one action, the browser runs it, and the model is called again with everything so far. Every click costs a network round trip and a prompt that has grown since the last one. I wanted to know what happens if the model writes the whole plan once and the steps are carried out locally.

So I built both, and a benchmark to compare them fairly. The full write-up is the preprint Plan Once, Ground Locally; this note is the short version, with the parts I would tell another engineer first.

1The setup

  • 200 tasks on seven public websites, each with ground truth scraped from the site itself or produced by a scripted browser run, and graded by a fixed regular expression. No model grades another model.
  • Four configurations, three repeats each, all on DeepSeek v4.1 flash at temperature 0: ReAct; ReAct with the same reading window and turn cap as the plan agent; plan-and-execute with Laya; and plan-and-execute with plain word overlap in place of Laya.
  • Laya is a ModernBERT-large decision model used zero-shot as a small “System 1”. It maps each planned step to a page element, judges when asynchronous content has arrived, and checks actions that changed nothing visible. It runs on the device, at a median of 31 ms per call.
  • 2,800 recorded runs: 2,400 on the main model and 400 more on a second, cheaper one. Every prompt, completion, tool result and model decision is in the published record.

2Accuracy is a tie once budgets match

Against the budget-matched ReAct baseline, planning once is not significantly different on success: 88.8% against 92.0%, a paired difference of −3.2 points with a 95% interval from −7.7 to +1.2.

The first comparison I ran looked much better, and it was wrong. Against plain ReAct the plan agent appeared 8.2 points more accurate (p < 0.001), but the whole gap came from plain ReAct’s smaller reading window and turn cap. Raising both gained ReAct 11.3 points. If your baseline is starved, your method will look brilliant.

3The savings are the real result

Table 1: Means per task over 200 tasks and three repeats, DeepSeek v4.1 flash.

ConfigurationSuccessTimeLLM callsInput tokens
ReAct80.7%32.8 s6.0320,523
ReAct, budget-matched92.0%42.2 s6.2224,697
Plan + Laya88.8%10.2 s2.034,060
Plan + word overlap86.5%10.8 s2.124,305

Against the budget-matched baseline, the plan agent used 83.6% fewer input tokens, 67.3% fewer LLM calls and 80.4% less money, and finished 4.1 times faster. On the 505 paired runs that both solved, it was 3,905 against 13,298 tokens, 1.88 against 4.41 calls, and 9.6 against 17.4 seconds.

4The clever part mattered less than I hoped

Laya was not significantly better than word overlap overall: +2.3 points, with an interval from −0.2 to +5.0 (p = 0.07). It earns its place where several controls look identical and context decides, as on saucedemo, where it solved 60 of 60 runs against 49 of 60. Elsewhere, matching words got most of the way.

5Where planning ahead breaks

On TodoMVC the plan configurations solved 0 and 2 of 30 runs, against 26 of 30 for budget-matched ReAct. A plan written before the page exists cannot name list items that only appear after you type them. Pages that build themselves as you act need the model back in the loop.

The cheaper second model reproduced the token and call savings and widened the accuracy gap in the plan agent’s favour, but not the money saving: in 7.5% of its runs it looped while writing the plan and hit the provider’s generation cap.

6What I would copy

  • Default to one plan, and call the model again only to replan after a failure and to write the answer: about two calls per task instead of six.
  • Ground steps locally with something cheap, and check that the clever version beats word overlap before paying for it.
  • Budget-match your baselines. A starved ReAct makes any alternative look good.
  • Keep a fallback for pages that only exist after you act on them.
  • Publish the traces, so every number can be recomputed from the recorded runs.

The paper, the 200 tasks with their ground truth, all 2,800 graded runs and the analysis scripts are on Zenodo, and the agent code is on GitHub.