# Plan Once, Ground Locally: Plan-and-Execute Web Agents Match ReAct Accuracy at One-Sixth of the Token Cost

Divyansh Shukla · Preprint · Zenodo · 2026-09-23 · DOI [10.5281/zenodo.22904595](https://doi.org/10.5281/zenodo.22904595) · CC BY 4.0

## Abstract

LLM web agents usually follow the ReAct pattern: the model is called again after every browser action, so each click is paid for with a network round trip and with a prompt that has grown since the last one. We study the alternative in which the model writes the whole plan in one call and the steps are executed locally, with a small on-device decision model (Laya, a ModernBERT-large “System-1” model used zero-shot) mapping each step to a page element, judging when asynchronous content has arrived, and checking actions that left the page unchanged.

We evaluate on a new benchmark of 200 tasks across seven public websites whose ground truth is scraped from the sites themselves or produced by scripted browser runs, and graded by fixed regular expressions. Every task was run three times under four configurations with the same LLM (DeepSeek v4.1 flash): ReAct, ReAct with a reading and turn budget matched to the plan agent, plan-and-execute with Laya, and an ablation that replaces Laya with word-overlap matching, for 2,400 runs, plus 400 further runs with a second model.

Against the budget-matched baseline, planning once is not significantly different on success (88.8% vs 92.0%; paired difference −3.2 points, 95% CI −7.7 to +1.2) while using 83.6% fewer input tokens, 67.3% fewer LLM calls and 80.4% less money, and finishing 4.1× faster.

Two negative results accompany that headline. Against the unmatched baseline the plan agent also looks more accurate (+8.2 points, p < 0.001), but the entire gap is an artefact of the baseline’s smaller reading window and turn cap. And the neural executor is not significantly better than word-overlap matching (+2.3 points, 95% CI −0.2 to +5.0). A second, cheaper model reproduces the token and call savings and widens the accuracy gap in the plan agent’s favour, but not the money saving: in 7.5% of its runs it looped while writing the plan and hit the provider’s generation cap.

## Results

200 tasks × 3 repeats per configuration, DeepSeek v4.1 flash at temperature 0. Means per task.

| Configuration | Success (95% CI) | Time | LLM calls | Input tokens | Cost |
|---|---|---|---|---|---|
| ReAct | 80.7% (77.3–83.6) | 32.8 s | 6.03 | 20,523 | $0.00186 |
| ReAct, budget-matched | 92.0% (89.6–93.9) | 42.2 s | 6.22 | 24,697 | $0.00202 |
| Plan + Laya | 88.8% (86.1–91.1) | 10.2 s | 2.03 | 4,060 | $0.00040 |
| Plan + word overlap | 86.5% (83.5–89.0) | 10.8 s | 2.12 | 4,305 | $0.00042 |

## Findings

- **Accuracy is a tie once budgets match.** Plan + Laya vs budget-matched ReAct: −3.2 points [−7.7, +1.2], not significant in any repeat (McNemar p = 0.31, 0.54, 0.12). The apparent win over plain ReAct comes from its smaller extract window and turn cap; raising both gains ReAct 11.3 points.
- **Efficiency is the real result.** 83.6% fewer input tokens, 67.3% fewer LLM calls, 80.4% less money and 4.1× less wall-clock time (Wilcoxon p < 1e-25 each). On the 505 paired runs both solved: 3,905 vs 13,298 tokens, 1.88 vs 4.41 calls, 9.6 vs 17.4 s.
- **The on-device grounder is not a significant win.** Laya vs word matching: +2.3 points [−0.2, +5.0], p = 0.07. It helps where identical controls need context (saucedemo: 60/60 vs 49/60 runs). Grounding calls take 31 ms median.
- **Where planning ahead fails.** TodoMVC: 0/30 and 2/30 runs for the plan configurations vs 26/30 for budget-matched ReAct. A plan written before the page exists cannot name list items that appear only after typing them.

## What the record contains

- The 200 tasks with ground truth and sources
- One graded row per run for all 2,800 runs
- Full per-run traces: every prompt, completion, tool result and model decision
- The agent code and the analysis scripts that recompute every number, table and figure

## Cite

```bibtex
@misc{shukla2026planonce,
  title     = {Plan Once, Ground Locally: Plan-and-Execute Web Agents
               Match ReAct Accuracy at One-Sixth of the Token Cost},
  author    = {Shukla, Divyansh},
  year      = {2026},
  month     = sep,
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.22904595},
  url       = {https://doi.org/10.5281/zenodo.22904595},
  note      = {Preprint}
}
```

## Links

- PDF: https://zenodo.org/records/22904595/files/plan-once-ground-locally.pdf
- Zenodo record: https://zenodo.org/records/22904595
- Code and data: https://github.com/TheDivyanshShukla/plan-once-ground-locally
