WEBMCP.COM

WindTunnel. An open benchmark for WebMCP.

WindTunnel compares WebMCP with other ways browser agents interact with websites.

Tasks solved
49/49
solved by every WebMCP configuration
+11.4% vs the median screen-driving agent (44 of 49)
Faster
2.5–7.5×
task completion
vs the median of other methods
Less cost
3–47×
per task
vs the median of other methods
Final score
+27–50%
higher final score
across WebMCP setups, vs the median of other methods

Leaderboard

Rank Configuration Interface Final score Tasks solved Attempt success Median costper task Median tokensper task Median timeper task
1Jev + Mercury 2.5New🏆WebMCP96.549/49141/147 (95.9%)$0.00119,7933.2s
2GPT-5.6 Luna · nativeWebMCP91.549/49146/147 (99.3%)$0.00242,5965.7s
3Gemini 3.6 Flash · nativeWebMCP88.249/49147/147 (100.0%)$0.00444,4537.2s
4Gemini 3.6 Flash · Stagehand v4WebMCP87.049/49146/147 (99.3%)$0.00444,3718.0s
5Sonnet 5 · nativeWebMCP86.049/49147/147 (100.0%)$0.00915,1726.8s
6Sonnet 5 · Stagehand v4WebMCP84.749/49147/147 (100.0%)$0.00955,1618.1s
7GPT-6 Astra · nativeNewWebMCP84.049/49146/147 (99.3%)$0.01712,5756.3s
8GPT-5.6 SOL · nativeWebMCP82.049/49145/147 (98.6%)$0.01212,5739.3s
9Claude Opus 5 · nativeWebMCP82.049/49147/147 (100.0%)$0.01404,7709.8s
10GPT-5.6 LunaComputer use71.345/49134/147 (91.2%)$0.017420,91418.3s
11GPT-6 AstraNewCode execution70.849/49147/147 (100.0%)$0.118710,98216.4s
12Jev + Mercury 2.5NewDOM (ultrafast)67.325/4976/147 (51.7%)$0.000813,8925.4s
13GPT-5.6 LunaDOM + vision67.043/49130/147 (88.4%)$0.032829,56119.8s
14GPT-5.6 LunaAccessibility tree65.840/49119/147 (81.0%)$0.019518,51716.0s
15Gemini 3.6 FlashComputer use64.843/49130/147 (88.4%)$0.020223,85733.7s
16GPT-5.6 SOLComputer use64.146/49134/147 (91.2%)$0.062716,23527.3s
17Sonnet 5DOM + vision63.948/49145/147 (98.6%)$0.210264,42429.3s
18GPT-6 AstraNewComputer use61.545/49135/147 (91.8%)$0.260820,56020.8s
19Sonnet 5Accessibility tree61.042/49128/147 (87.1%)$0.037810,76237.5s
20Claude Opus 5Computer use56.945/49134/147 (91.2%)$0.138947,14150.4s
21Sonnet 5Computer use56.539/49119/147 (81.0%)$0.069757,70131.7s

Final score: The final score combines attempt success rate (60%), median cost per task (20%), and median agent time per task (20%). Cost and time are log-normalized (large differences are compressed so extreme values do not dominate); higher is better. Tokens are shown separately and are not scored twice.

Methodology

We operate eight real websites across a wide range of tasks — 49 tasks, three attempts each. There are four ways to operate the website:

WEBMCP

Tool calling

The website exposes direct actions for the agent to call.

SCREENSHOTS

Computer use

The agent reads rendered images of the page and acts by coordinate.

PAGE STRUCTURE

DOM

The agent reads the page's DOM and accessibility tree.

CODE EXECUTION

Browser code

The model writes code that the agent executes to read and drive the page.

Everything is fully reproducible — the methodology, code, task definitions, and full run transcripts are on github.com/nekuda-ai/WindTunnel.

Changelog

v1.2

Sep 18, 2026 Latest
  • Added Jev + Mercury 2.5 with WebMCP and DOM controls without WebMCP.
  • Jev + Mercury 2.5 with WebMCP leads the final score. The board now includes 21 configurations.

v1.1

Sep 6, 2026
  • Added the complete_checkout WebMCP tool, fixing the checkout task on the online store.
  • Re-measured all eight WebMCP configurations. Checkout improved from 0/3 → 3/3 for each.
  • Added GPT-6 Astra benchmarks for WebMCP, screenshots, and code execution.

v1.0

Aug 20, 2026
  • Initial benchmark release: 16 configurations, 2,352 attempts.

View full changelog →