WebsiteBench

WebsiteBench : Can AI agents rebuild websites through browser exploration?

Try the task Cite Paper, code and data coming soon

74 expert-built websites50,356 hidden tests9 models16 hours per run

Reconstruction performance and mean resource use per task for nine models on 74 websites.
# Model
1Claude Opus 5.5Anthropic, frontier model22.563.518.034.8$261.1930.5M2,107.1
2Claude Opus 5Anthropic, frontier model18.743.114.033.2$221.7429.2M980.5
3GPT 5.6 SolOpenAI, frontier model15.638.210.137.4$11520.5M212.3
4GPT 5.6 LunaOpenAI, frontier model13.229.88.930.0$3.814.3M174.0
5Grok 4.6xAI, frontier model11.637.87.829.6$76.938.0M403.9
6GLM 5.3 FlashZ.ai, open-source model9.720.46.623.1$2.437.0M318.8
7Kimi K3Kimi, open-source model8.922.64.727.1$37.291.6M604.3
8DeepSeek V4.1 FlashDeepSeek, open-source model8.523.74.226.5$4.6167.8M572.6
9Qwen 3.8 MaxQwen Team, open-source model5.722.61.820.2$9.538.3M252.1

Every frontier model beats every open-source model, overall and in each test group; a heavier rule separates the two groups. Scores are percentages of hidden tests passed. Cost, tokens and turns are means per task. Bold marks the best value in each column, and the tick in each Overall bar marks 25%. Under each model, R, B and S give its Request, Browser and System pass rates.

Why even the best agent passes fewer than one test in four

/why
1Quits flows before the end 1~10 paths tried per site
1UnderexplorationHumans explore 28 pages, agents about 10 paths
2Right labels, wrong images 2?tab=coming-soon ignored 2Booking routes return 404 2Review fails after restart
2Failure to implement observed behavior1 of 111 observed behaviors implemented as observed
3Stops at 1.0 h of 16 with 0% 3Declares done, bugs remain
3Early stoppingMedian run ends at 3.57 of 16 hours
WebsiteBench 2,107 average steps per site for the best model and still only 22.5% of hidden tests pass
The three failure modes the paper identifies, with examples drawn from its analyses and case studies.

The agent can use the website but never read its code

/task

Each task starts from a local website that an expert built to reproduce selected features of a widely used service. The agent reaches it only through a browser: it can navigate, click, type and take screenshots, but never read the HTML, scripts or server code. It has up to 16 hours to write a full-stack site that looks and behaves the same, and it can return to the reference as often as it likes while it codes and tests. Hidden tests then compare the two sites. The agent never sees those tests.

Illustrative mock, not a benchmark taskIllustrative mock 2 of 6 hidden tests passed
Hidden tests
  • Browser L1Passed. Header and navigation match the referencesimilarity 0.93 ≥ 0.75
  • Browser L1Failed. Category tiles match the referencesimilarity 0.41 < 0.75
  • RequestFailed. GET /shop?tab=deals opens the Deals tabopened All
  • Browser L2Failed. Book grooming, then see the confirmation404 /grooming/book
  • Browser L3Failed. Cart items survive a server restart2 items became 0
  • SystemPassed. Builds offline and answers /healthz{"status":"ok"}
Drag the divider to compare a mock reference site with a mock rebuild. Click around, restart the servers or run the tests. Each defect mirrors a failure reported in the paper. Real tasks average about 680 hidden tests each.

What the agent can use

  • Screenshots and interface state of the reference
  • Navigation, clicks, typing, scrolling and form submissions
  • Images and other media files from the reference
  • Its own email inbox for sign-up and verification flows
  • The task instructions and the runtime contract
  • A shell, build tools and a local browser for its own site

What the agent never sees

  • The reference’s HTML, JavaScript and CSS
  • API responses and arbitrary scripts run in the page
  • The server’s source code
  • The hidden tests, reference captures and scores

Three ways the rebuilds go wrong

/findings

We traced 20 Grok 4.6 runs from what the agent observed, to the code it wrote, to the tests it failed, and compared runs with and without a human exploration trace.

1Underexploration

Agents visit the pages but seldom follow a workflow to its end

Agents often stop a multi-step interaction before its final step, so they never see what a checkout confirms or what a saved record looks like afterwards. Experts exploring the same sites reached about ten times as many paths in a twentieth of the time.

Giving the agent a human exploration trace, a log of browser actions with no page content, raised the mean score on 18 paired websites from 2.8% to 9.8%. The trace helped on 10 sites, tied on 6 and hurt on 2 (sign test p = 0.039).

Human and agent exploration compared
TimeActionsReach
Human expertmedian just over 5 min450+ actions28 pages
Agent1.5 to 2.5 h200 to 300 stepsabout 10 paths

Mean score on 18 paired websites

Agent explores alone2.8%
Agent with a human trace9.8%
10 higher6 equal2 lower

2Failure to implement observed behavior

What agents observe rarely reaches their code

Across 20 websites we found 111 behaviors that the agent had observed during exploration and that its recorded trajectory and submitted code let us check: page contents, workflows and persisted state. One was implemented the way the agent had seen it. The other 110 did not match the observation.

Effort is lopsided too. More than half of all tool calls explored the reference, and only one in ten edited code.

111 observed requirements

1 implemented as observed110 mismatched

13,359 tool calls in 20 runs

Exploring the reference 54.8% Editing code 10.0% Checking the candidate 27.0% Other, such as setup 8.2%

3Early stopping

Agents stop with most of the clock left

Runs may take 16 hours, yet the median run ended after 3.57 hours. We then enabled Claude Code’s /goal hook for Kimi K3 on the Lowe’s task, so that a review had to accept each attempt to stop.

At 10.25 hours the agent declared its repairs complete. The hook rejected the exit, and the agent went on to check that cart, order, account and session state survived a restart. The score rose from 0% to 31.3%.

Run length on a 16-hour budget

Median run3.57 h
Kimi K3 on Lowe’s1.0 h, score 0%
Same task with /goal10.9 h, score 31.3%

At 10.25 h the hook rejected the agent’s first attempt to stop.

Hours
1.0 → 10.9
Steps
130 → 879
Tokens
10.6M → 87.9M
Score
0% → 31.3%

74 expert-built sites, 50,356 tests, no partial credit

/tests

74 websites across nine categories

Sixteen web developers designed and built the tasks under shared guidelines. For each website, one expert implements the reference and another reviews it for correctness and workflow coverage, then writes the hidden tests. Each site reproduces selected pages and workflows of a widely used service, with email and payments replaced by local services.

Sites named in the paper reproduce features of

  • AMC Theatres
  • Backcountry
  • Fandango
  • L.L.Bean
  • Lowe’s
  • Medium
  • Meetup
  • Petfinder
  • PetSmart
  • Samsung
  • Target
  • Tripadvisor

Three groups of hidden tests

T1

Request

Direct HTTP checks of page responses and native form submissions.

T2

Browser

Scripted journeys through the real interface, at three levels.

  1. L1An immediate response on one page, such as input validation.
  2. L2A state change read back in a later observation, on the same page or another, such as an edited record on its detail page.
  3. L3Persistence, other users and recovery, such as data after a restart.

Some Browser tests add visual checkpoints that compare screenshots of marked regions with the reference.

T3

System

Runtime requirements: offline builds, restarts, process health and security.

All or nothing, test by test

  • A test earns its point only if its functional checks pass and every visual checkpoint reaches the threshold τ = 0.75.
  • Journey tests must pass under two independent executors, Playwright and Browser Use.
  • Visual similarity is SSIM measured against a flat-color baseline, so large blank areas count for much less.
  • Every test weighs the same and there is no partial credit. A submission that copies the reference scores zero.

Cite WebsiteBench

/cite
@misc{websitebench2026,
  title  = {WebsiteBench: Can AI Agents
            Rebuild Websites through
            Browser Exploration?},
  author = {Anonymous},
  year   = {2026}
}