WEBMCP.COM ← Back to the directory

Implementation audit methodology

How we audit WebMCP implementation

01 What the audit checks

Your WebMCP implementation grade summarizes the tool definitions and registration details reviewed in your audit. The report shows what we inspected, the issues we found, and suggested fixes.

Basic setup
Live scans and private definition reviews use the same factual checks for tool names, descriptions, complete input and output schemas, and recorded registration or handler errors. Invalid nested schema rules count as defects too.
Clear definitions
An AI reviewer checks whether each tool clearly explains what it does, which inputs it accepts, and what it returns. Each finding points to the relevant tool definition and suggests a specific fix.
Testing behavior
Checking how tools work requires running them in your app. Authentication, input validation, error handling, page updates, and complete user flows need separate execution tests.

The scan checks your starting page and up to five additional pages. If only one page is checked, pages fail to load, or page suggestions cannot all be checked, the report shows a limited coverage notice. Its grade reflects the pages actually checked, not a complete inventory of your website.

Help the scanner find your important pages

Add an optional pages array to your existing /.well-known/webmcp.json, such as "pages": ["/products", "/products/example"]. Put a representative page of each type first. These pages must be on the same origin as your homepage after redirects.

This hint is a webmcp.com extension, not a WebMCP standard field. The scanner checks these suggestions first, then pages found in previous scans, sitemaps and homepage links. It still inspects each page for live tools.

02 How we determine your grade

We grade the issues found during the audit, then use the lowest grade as the overall result. For example, a missing tool description gives a B. Fixing it lets the next audit reflect the remaining findings.

We review both tools registered in JavaScript and tools defined through HTML forms. We look for clear descriptions and inputs that match the tool's purpose. Tools that take no inputs can omit an input schema, and output schemas are optional.

Our review follows Chrome's WebMCP best practices: explain actions and results clearly, define meaningful input types, and validate inputs in your application. WebMCP.com defines the grades below.

03 What each grade means

GradeWhat the audit found
A+The reviewed definitions and registration details passed the audit with no suggested changes.
ASmall, optional improvements to the tool definitions.
A-A tool's description or expected result needs clarification, or its purpose overlaps with another tool.
B+Inputs need clearer types, units, or accepted values, or require avoidable conversions.
BA tool is missing its description, or gives unclear or conflicting instructions about what it does.
B-The scan captured an invalid schema or an invalid function for executing a tool.
CA tool is missing its name, or the browser rejected its registration.

Use the report to fix the issues found, then request another audit. Live definition grades currently use webmcp-implementation-v4. A fresh audit applies the current grading rules; older grades are not reused as current results.

04 Where your grade comes from

Not graded
This means we still need enough information to assign a grade. The report explains why, such as blocked access, no tools detected on the inspected pages, or an unfinished review.
Before deployment
A private definition review assesses the definitions you submit. A Code review grade instead assesses the selected implementation files under its separate code-review rubric. Neither private review updates your public grade or establishes that the code is deployed.
After deployment
Your Live definition grade comes from a scan of your website. The report includes the scan date, inspected pages, and reviewed definitions. After deploying changes, run a fresh scan. The grade appears in your report, not in the public directory; directory tool updates still go through review.
Uncertain evidence
If a supplied schema cannot be fully validated, or different definitions share a name that the reviewer cannot distinguish, the report explains the problem without issuing a grade. Identical definitions repeated across pages count once.
Testing tool behavior
Run separate tests to see how agents choose and use your tools. Google's WebMCP Evals CLI supports tool-selection evaluations and tests of specific tool calls. These results describe the cases you tested. See Chrome's evaluation guidance.

05 Tool categories

Every tool is sorted into one of three categories — a trust ladder for how freely an agent can call it. The category is assigned automatically by classifying each tool's name, description, and input schema during the scan, so it's inferred, not declared by the site.

Answer
Read-only tools that return information and leave the page untouched: search, details, availability, policies. No side effects; safe for an agent to call freely.
Action
Tools that change state or drive the page on the user's behalf without commitment: carts, filters, forms, navigation, redirects. Reversible.
Sensitive Action
Tools that involve money or commitment: booking, checkout, ordering, subscribing. The highest trust bar — agents should require explicit user confirmation.