v1.2
Sep 18, 2026 Latest- Added Jev + Mercury 2.5 with WebMCP and DOM controls without WebMCP.
- Jev + Mercury 2.5 with WebMCP leads the final score. The board now includes 21 configurations.
WindTunnel compares WebMCP with other ways browser agents interact with websites.
| Rank | Configuration | Interface | Final score | Tasks solved | Attempt success | Median costper task | Median tokensper task | Median timeper task |
|---|---|---|---|---|---|---|---|---|
| 1 | Jev + Mercury 2.5New🏆 | WebMCP | 96.5 | 49/49 | 141/147 (95.9%) | $0.0011 | 9,793 | 3.2s |
| 2 | GPT-5.6 Luna · native | WebMCP | 91.5 | 49/49 | 146/147 (99.3%) | $0.0024 | 2,596 | 5.7s |
| 3 | Gemini 3.6 Flash · native | WebMCP | 88.2 | 49/49 | 147/147 (100.0%) | $0.0044 | 4,453 | 7.2s |
| 4 | Gemini 3.6 Flash · Stagehand v4 | WebMCP | 87.0 | 49/49 | 146/147 (99.3%) | $0.0044 | 4,371 | 8.0s |
| 5 | Sonnet 5 · native | WebMCP | 86.0 | 49/49 | 147/147 (100.0%) | $0.0091 | 5,172 | 6.8s |
| 6 | Sonnet 5 · Stagehand v4 | WebMCP | 84.7 | 49/49 | 147/147 (100.0%) | $0.0095 | 5,161 | 8.1s |
| 7 | GPT-6 Astra · nativeNew | WebMCP | 84.0 | 49/49 | 146/147 (99.3%) | $0.0171 | 2,575 | 6.3s |
| 8 | GPT-5.6 SOL · native | WebMCP | 82.0 | 49/49 | 145/147 (98.6%) | $0.0121 | 2,573 | 9.3s |
| 9 | Claude Opus 5 · native | WebMCP | 82.0 | 49/49 | 147/147 (100.0%) | $0.0140 | 4,770 | 9.8s |
| 10 | GPT-5.6 Luna | Computer use | 71.3 | 45/49 | 134/147 (91.2%) | $0.0174 | 20,914 | 18.3s |
| 11 | GPT-6 AstraNew | Code execution | 70.8 | 49/49 | 147/147 (100.0%) | $0.1187 | 10,982 | 16.4s |
| 12 | Jev + Mercury 2.5New | DOM (ultrafast) | 67.3 | 25/49 | 76/147 (51.7%) | $0.0008 | 13,892 | 5.4s |
| 13 | GPT-5.6 Luna | DOM + vision | 67.0 | 43/49 | 130/147 (88.4%) | $0.0328 | 29,561 | 19.8s |
| 14 | GPT-5.6 Luna | Accessibility tree | 65.8 | 40/49 | 119/147 (81.0%) | $0.0195 | 18,517 | 16.0s |
| 15 | Gemini 3.6 Flash | Computer use | 64.8 | 43/49 | 130/147 (88.4%) | $0.0202 | 23,857 | 33.7s |
| 16 | GPT-5.6 SOL | Computer use | 64.1 | 46/49 | 134/147 (91.2%) | $0.0627 | 16,235 | 27.3s |
| 17 | Sonnet 5 | DOM + vision | 63.9 | 48/49 | 145/147 (98.6%) | $0.2102 | 64,424 | 29.3s |
| 18 | GPT-6 AstraNew | Computer use | 61.5 | 45/49 | 135/147 (91.8%) | $0.2608 | 20,560 | 20.8s |
| 19 | Sonnet 5 | Accessibility tree | 61.0 | 42/49 | 128/147 (87.1%) | $0.0378 | 10,762 | 37.5s |
| 20 | Claude Opus 5 | Computer use | 56.9 | 45/49 | 134/147 (91.2%) | $0.1389 | 47,141 | 50.4s |
| 21 | Sonnet 5 | Computer use | 56.5 | 39/49 | 119/147 (81.0%) | $0.0697 | 57,701 | 31.7s |
Final score: The final score combines attempt success rate (60%), median cost per task (20%), and median agent time per task (20%). Cost and time are log-normalized (large differences are compressed so extreme values do not dominate); higher is better. Tokens are shown separately and are not scored twice.
We operate eight real websites across a wide range of tasks — 49 tasks, three attempts each. There are four ways to operate the website:
The website exposes direct actions for the agent to call.
The agent reads rendered images of the page and acts by coordinate.
The agent reads the page's DOM and accessibility tree.
The model writes code that the agent executes to read and drive the page.
Everything is fully reproducible — the methodology, code, task definitions, and full run transcripts are on github.com/nekuda-ai/WindTunnel.
complete_checkout WebMCP tool, fixing the checkout task on the online store.