divyanshshukla.com
Preprint · v1.0.0

Plan Once, Ground Locally: Plan-and-Execute Web Agents Match ReAct Accuracy at One-Sixth of the Token Cost

Divyansh Shukla

Preprint, Zenodo · doi:10.5281/zenodo.22904595 · CC BY 4.0

23 September 2026

PDFZenodo recordCode and data

Abstract

LLM web agents usually follow the ReAct pattern: the model is called again after every browser action, so each click is paid for with a network round trip and with a prompt that has grown since the last one. We study the alternative in which the model writes the whole plan in one call and the steps are executed locally, with a small on-device decision model (Laya, a ModernBERT-large “System-1” model used zero-shot) mapping each step to a page element, judging when asynchronous content has arrived, and checking actions that left the page unchanged.

We evaluate on a new benchmark of 200 tasks across seven public websites whose ground truth is scraped from the sites themselves or produced by scripted browser runs, and graded by fixed regular expressions. Every task was run three times under four configurations with the same LLM (DeepSeek v4.1 flash): ReAct, ReAct with a reading and turn budget matched to the plan agent, plan-and-execute with Laya, and an ablation that replaces Laya with word-overlap matching, for 2,400 runs, plus 400 further runs with a second model.

Against the budget-matched baseline, planning once is not significantly different on success (88.8% vs 92.0%; paired difference −3.2 points, 95% CI −7.7 to +1.2) while using 83.6% fewer input tokens, 67.3% fewer LLM calls and 80.4% less money, and finishing 4.1× faster.

Two negative results accompany that headline. Against the unmatched baseline the plan agent also looks more accurate (+8.2 points, p < 0.001), but the entire gap is an artefact of the baseline’s smaller reading window and turn cap. And the neural executor is not significantly better than word-overlap matching (+2.3 points, 95% CI −0.2 to +5.0). A second, cheaper model reproduces the token and call savings and widens the accuracy gap in the plan agent’s favour, but not the money saving: in 7.5% of its runs it looped while writing the plan and hit the provider’s generation cap.

1Results

Each of the 200 tasks has ground truth scraped from the site itself or produced by a scripted browser run, and is graded by a fixed regular expression, so no model judges another. The budget-matched ReAct baseline gets the same reading window and turn cap as the plan agent; it is the fair comparison.

Table 1: 200 tasks on seven public websites, three repeats per configuration, DeepSeek v4.1 flash at temperature 0. Success is the pooled rate with its 95% interval; the other columns are means per task. Best in bold.

ConfigurationSuccessTimeLLM callsInput tokensCost
ReAct4,000-char extract, 15 turns80.7%77.3–83.632.8 s6.0320,523$0.00186
ReAct, budget-matched12k-char extract, 30 turns92.0%89.6–93.942.2 s6.2224,697$0.00202
Plan + Layaone plan, grounded on device88.8%86.1–91.110.2 s2.034,060$0.00040
Plan + word overlapablation: no Laya86.5%83.5–89.010.8 s2.124,305$0.00042
Figure 1: Mean input tokens per task. Writing the plan once cuts input tokens by 83.6% against the budget-matched ReAct baseline, at no significant cost in success (Table 1).

2Findings

Accuracy is a tie once budgets match. Plan + Laya vs budget-matched ReAct: −3.2 points [−7.7, +1.2], not significant in any repeat (McNemar p = 0.31, 0.54, 0.12). The apparent win over plain ReAct comes from its smaller extract window and turn cap; raising both gains ReAct 11.3 points.

Efficiency is the real result. 83.6% fewer input tokens, 67.3% fewer LLM calls, 80.4% less money and 4.1× less wall-clock time (Wilcoxon p < 1e-25 each). On the 505 paired runs both solved: 3,905 vs 13,298 tokens, 1.88 vs 4.41 calls, 9.6 vs 17.4 s.

The on-device grounder is not a significant win. Laya vs word matching: +2.3 points [−0.2, +5.0], p = 0.07. It helps where identical controls need context (saucedemo: 60/60 vs 49/60 runs). Grounding calls take 31 ms median.

Where planning ahead fails. TodoMVC: 0/30 and 2/30 runs for the plan configurations vs 26/30 for budget-matched ReAct. A plan written before the page exists cannot name list items that appear only after typing them.

3What the record contains

  • The 200 tasks with ground truth and sources
  • One graded row per run for all 2,800 runs
  • Full per-run traces: every prompt, completion, tool result and model decision
  • The agent code and the analysis scripts that recompute every number, table and figure

4Cite

BibTeX
@misc{shukla2026planonce,
  title     = {Plan Once, Ground Locally: Plan-and-Execute Web Agents
               Match ReAct Accuracy at One-Sixth of the Token Cost},
  author    = {Shukla, Divyansh},
  year      = {2026},
  month     = sep,
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.22904595},
  url       = {https://doi.org/10.5281/zenodo.22904595},
  note      = {Preprint}
}