Crawl Readiness MCP Server
Crawl Readiness MCP Server gives your AI assistant eight AI-SEO tools to audit whether AI crawlers can read a site and then generate the files to fix it.
Audit (no API key):
check_ai_readinessscores a site 0–100 against 50+ AI crawlers with a prioritized fix list;validate_schemaaudits JSON-LD (20+ types);validate_robotsaudits robots.txt line-by-line plus AI-bot coverage;check_content_paritycompares human vs AI-crawler views (word overlap, cloaking, JS-only shells)Generate (free API key):
generate_llms_txt,generate_robots_txt(presets: allow-all, search-only, recommended, block-all), andgenerate_schema(Organization/WebSite/Article JSON-LD plus BreadcrumbList/FAQPage templates)Monitor (read-only, API key):
get_monitor_trendshows whether ChatGPT, Claude, Perplexity, and Google AI mention your brand, with share of voice, per-provider breakdown, and competitor trends — call with no argument to list all monitored brandsWorkflow: chain tools to audit, fix, and write files straight into your project without copy-pasting
Audits whether a website is accessible to Google AI crawlers, including robots.txt status and structured-data issues.
Audits whether a website is accessible to Perplexity's AI crawler, including robots.txt status and content parity.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Crawl Readiness MCP ServerIs my site AI-readable? URL: mydomain.com"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Crawl Readiness MCP Server
AI SEO tools inside Claude, Cursor, and other MCP-compatible clients.
Ask Claude "is my site AI-readable?" and it'll actually check — running a live audit against 50+ AI crawlers (ChatGPT, Claude, Perplexity, Google AI, and more). Then ask it to fix what's broken, and it can generate a proper llms.txt, robots.txt, or JSON-LD schema and write it straight into your project.
No more copy-paste between your editor and yet another SEO tool.
What it does
Eight tools your AI assistant can call:
Audit (no signup):
check_ai_readiness— Score a site 0–100 against 50+ AI crawlers, with a prioritized fix list; extras that don't affect the score are marked as suchvalidate_schema— Audit JSON-LD structured data on any URL (20+ types, per-type rules)validate_robots— Line-by-line robots.txt auditcheck_content_parity— Compare what humans see vs what AI crawlers see
Generate (requires free API key):
generate_llms_txt— Auto-crawl a site and produce a properly formatted llms.txtgenerate_robots_txt— AI-crawler-aware robots.txt with preset policiesgenerate_schema— Organization / WebSite / Article JSON-LD ready to paste
Monitor (read-only, requires free API key):
get_monitor_trend— See whether ChatGPT, Claude, Perplexity & Google AI mention your brand vs competitors, and how that's trending week over week (reads your LLM Monitor projects; doesn't trigger runs)
Related MCP server: SEO MCP Server
Install
For Claude Desktop
Open your config file:
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%\Claude\claude_desktop_config.json
Add this to the mcpServers block:
{
"mcpServers": {
"crawl-readiness": {
"command": "npx",
"args": ["-y", "crawl-readiness-mcp"]
}
}
}Restart Claude Desktop. You'll see "crawl-readiness" listed in the tools panel.
For Cursor
Open Cursor Settings → Features → Model Context Protocol → Add new global MCP server.
{
"crawl-readiness": {
"command": "npx",
"args": ["-y", "crawl-readiness-mcp"]
}
}For Cline / Windsurf / Continue
Same shape — add crawl-readiness to your MCP servers list with npx -y crawl-readiness-mcp as the command.
Unlock the generator tools
The four audit tools work out of the box, no signup. The three generator tools need an API key — free at crawlreadiness.com/dashboard (50 checks/month on the free tier, no credit card).
Once you have a key, add it via env in the same config:
{
"mcpServers": {
"crawl-readiness": {
"command": "npx",
"args": ["-y", "crawl-readiness-mcp"],
"env": {
"CRAWL_READINESS_API_KEY": "cr_live_xxxxxxxxxxxx"
}
}
}
}Restart the client. All eight tools are now callable.
Example prompts
Once installed, try any of these in your AI client:
"Is my portfolio site AI-readable? URL: yourdomain.com"
"Check if ChatGPT can see acme-inc.com"
"Validate the JSON-LD on stripe.com — is anything missing?"
"Compare what humans see vs what AI crawlers see on producthunt.com"
"Generate an llms.txt for my site and save it to
/public/llms.txt""My site scored 68/100 — walk me through the fixes"
Your AI can chain tools automatically — audit, identify issues, fix them, and write files to your project without you copy-pasting anything.
Configuration
Env var | Purpose | Default |
| Free API key for generator tools | none |
| Override the API base URL (for testing) |
|
What tools receive from the server
Each tool returns the full JSON response from the Crawl Readiness API — the same data the web dashboard renders. Your AI has access to:
check_ai_readiness: score, robots.txt status per crawler, structured-data flags, homepage signals, prioritized fix list (fixes marked
extra: trueare optional and not scored)validate_schema: per-block issues (errors, warnings, info), suggested missing types
validate_robots: line-by-line issues, AI-bot coverage summary
check_content_parity: verdict + per-crawler HTTP status, word overlap %, warnings
generate_llms_txt: the file content + companion robots.txt snippet
generate_robots_txt: merged robots.txt with the AI policy applied
generate_schema: JSON-LD scripts for each detected schema type
get_monitor_trend: per-brand mention rate, share of voice, per-provider breakdown, competitor comparison, trend over time, and short example answers
Support
Docs: crawlreadiness.com/mcp
Email: support@crawlreadiness.com
About
Crawl Readiness is a suite of AI-SEO tools that check whether AI systems can access your website and help you fix what's blocked. Every generator tool in this MCP server is also available on the web at crawlreadiness.com/tools. The MCP server just makes the same tools callable from inside your AI assistant, so the "audit → fix → apply" loop happens in one conversation without leaving your editor.
Releasing (maintainers)
npm run release validates server.json and prints the publish steps. The last step uses the mcp-publisher CLI, which is not part of this repository: download it from the MCP Registry releases and keep it on your PATH or in the repo root (it is git-ignored there).
License
MIT © 2026 Crawl Readiness / Tundrastone
Available Tools
8 toolscheck_ai_readinessCheck AI ReadinessARead-only
Check whether AI crawlers (ChatGPT, Claude, Perplexity, Google AI, and 50+ others) can access a website. Returns a 0-100 AI readiness score, per-crawler access status, detected AI-specific files (llms.txt, agents.json), structured data presence, meta signals, and a prioritized fix list. Use this as the first step in any AI SEO audit.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The website URL to check (e.g. 'example.com' or 'https://example.com/page'). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, and the description is consistent with these. It adds substantial behavioral context by enumerating exactly what checks are performed (AI crawler access, file detection, structured data, meta signals) and the output contents. This goes beyond the annotations, which only signal safety, not the scope of the audit. No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no redundancy. The first sentence front-loads the core action and return payload; the second provides usage guidance. Every word earns its place, and the structure naturally surfaces the important 'first step' instruction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema exists, the description fully enumerates the output components (score, per-crawler status, AI-specific files, structured data, meta signals, fix list), which is essential for the agent to understand what to expect. The one parameter is well-covered by the schema. The usage context is clear. Nothing an agent needs to invoke this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% – the schema fully describes the 'url' parameter. The description only refers to 'a website' without adding format or constraints beyond the schema (e.g., no mention of protocol handling or query string). Baseline 3 is correct when the schema carries the parameter documentation, and the description adds no extra semantic value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Check') and resource ('AI crawlers can access a website'), then enumerates the detailed return values (readiness score, per-crawler status, AI-specific files, structured data, meta signals, fix list). This clearly distinguishes it from sibling validators (validate_robots, validate_schema, etc.) as a comprehensive first-step audit, not a single-aspect check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the agent to 'Use this as the first step in any AI SEO audit', giving clear positional guidance among siblings. It does not explicitly list when NOT to use it (e.g., for deep validation of a single aspect, use validate_robots), but the first-step framing implies that granular validators follow. A minor gap, so a 4 is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_content_parityContent Parity CheckARead-only
Compare what human browsers see vs what AI crawlers see. Fetches a page four times in parallel — as Chrome, GPTBot, ClaudeBot, and PerplexityBot — and reports word-overlap %, title/description differences, and warnings about JS-only shells, cloaking, or edge-based bot blocking.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to check parity on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as read-only and open-world, so the safety profile is covered. The description adds substantial behavioral context: it performs four parallel fetches, specifies the exact user agents, and enumerates what it detects (JS-only shells, cloaking, edge-based bot blocking), which goes well beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the purpose and followed by the method and output details. Every phrase adds value and there is no filler or repetition of annotation fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With only one parameter, no output schema, and annotations already covering safety, the description is complete: it states what the tool does, how it fetches, and what reports/warnings the agent should expect. Nothing essential for calling and interpreting the tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single url parameter already has a clear description. The tool description does not add parameter-level details, which matches the baseline expectation when the schema fully documents parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific, meaningful comparison ('human browsers vs AI crawlers') and then names the concrete mechanism (four parallel fetches as Chrome, GPTBot, ClaudeBot, PerplexityBot) and outputs (word-overlap %, title/description differences, warnings). This clearly distinguishes it from siblings like check_ai_readiness or validate_robots by focusing on parity detection rather than readiness or validation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use case is implied clearly: call this when you need to compare rendering across browsers and AI crawler user agents. However, it never explicitly states when not to use it or which sibling alternative to prefer, leaving the selection decision to the agent based only on sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_llms_txtGenerate llms.txtA
Generate a properly formatted llms.txt file for a website. Crawls the site's sitemap, groups pages by section, pulls page titles and descriptions, and produces both the llms.txt content and a companion robots.txt snippet. The AI client can then write the returned content to disk in the user's project. Requires an API key.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The website to generate llms.txt for. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavioral context beyond annotations: it performs an external crawl (aligning with openWorldHint), requires an API key, and returns content for the AI client to write to disk rather than writing files itself. It does not detail failure modes or rate limits, but those are minor given the annotations already flag the open-world, non-read-only nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences carry the full purpose, process, outputs, client action, and prerequisite with no filler. The main purpose is front-loaded and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no output schema, the description explains what is produced (llms.txt content and a companion robots.txt snippet), how the client should handle the result, and the auth prerequisite. Minor gaps remain around exact return structure and API-key mechanism, but the core information needed for selection and invocation is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides 100% coverage for the single url parameter, so the baseline is appropriate. The description reinforces that the URL should be a website whose sitemap will be discovered and crawled, but it does not specify URL format, scheme requirements, or how the API key should be supplied. This is adequate but not additive beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and deliverable: generate a properly formatted llms.txt file. It adds concrete process details (crawls sitemap, groups pages by section, pulls titles and descriptions) and names the companion output of a robots.txt snippet, making it clearly distinct from sibling tools like generate_robots_txt or generate_schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It is clear that this tool is for generating llms.txt for a website, and the prerequisite of an API key is explicit. It does not directly contrast with siblings or state when not to use it, but the primary artifact is unambiguous and the description provides enough context for an agent to know when this tool applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_robots_txtGenerate robots.txtA
Generate a complete, ready-to-save AI-crawler-aware robots.txt for a website. Fetches the existing robots.txt (if any) and returns the finished file in generated.robotsTxt — the existing rules with an AI-crawler policy section merged in — plus the per-bot allow/block breakdown. Presets: 'allow-all' (public businesses), 'search-only' (allow AI search, block training), 'recommended' (allow major AI assistants that cite sources, block training-only bots), 'block-all'. Write generated.robotsTxt to the site's /robots.txt. Requires an API key.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The website to generate a robots.txt for. | |
| preset | No | The AI-crawler policy preset to apply. | recommended |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint=false annotation, the description discloses meaningful side effects: it fetches the existing robots.txt, writes generated.robotsTxt to /robots.txt, and requires an API key. It also describes the return payload (generated file plus per-bot allow/block breakdown), giving the agent a clear picture of behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but each sentence earns its place: purpose, workflow, presets, side effect, and auth requirement. The important behavioral information is front-loaded and the preset list is structured and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no output schema, the description covers the return value, the side effect, the presets, and the auth requirement. It could be more complete by limiting when to use it vs. sibling tools, but it is adequate for a competent agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real value by explaining the meaning of each preset ('allow-all', 'search-only', 'recommended', 'block-all') and their intended use cases. It does not add syntax details for url, but the schema already documents that parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific deliverable ('AI-crawler-aware robots.txt') and a concrete workflow (fetch, merge, write), which differentiates it from sibling validation/generation tools. The opening sentence states exactly what the tool produces for a website.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool: to generate a ready-to-save robots.txt with an AI-crawler policy, and it enumerates preset policies for different site stances. It does not explicitly say 'use validate_robots instead for validation', but this is not necessary for correct selection given the sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_schemaGenerate JSON-LD SchemaA
Generate JSON-LD structured data for a website. Detects the site's name, logo, social profiles, contact info, and article metadata, then produces Organization, WebSite, and (when applicable) Article schemas plus starter templates for BreadcrumbList and FAQPage. Returns each schema as a block ready to paste into the site's . Requires an API key.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The website to generate schema for. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond annotations, the description discloses meaningful behavioral traits: it detects site metadata, produces multiple schema types plus starter templates, returns ready-to-paste script blocks, and requires an API key. This aligns with readOnlyHint=false and openWorldHint=true while adding concrete expectations about the tool's operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four sentences, each carrying distinct information: the core action, the detected metadata, the output format, and the API key requirement. There is no filler or redundancy, and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no output schema, the description covers the essential calling context: what it detects, what schemas it produces, how output is returned, and the required API key. It does not explain API key acquisition or error behavior, but the agent has enough to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already fully describes the single 'url' parameter as 'The website to generate schema for,' giving 100% schema description coverage. The description reinforces that the URL points to the target website but adds no parameter-level detail, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Generate') and resource ('JSON-LD structured data for a website'), then enumerates concrete output types including Organization, WebSite, Article, BreadcrumbList, and FAQPage. This clearly differentiates it from siblings such as validate_schema or generate_robots_txt.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool: generating JSON-LD schema for a website, and it notes the API key prerequisite. It does not explicitly name alternative tools or state when not to use it, but the artifact type and sibling names make the intended use reasonably unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_monitor_trendGet LLM Monitor TrendARead-only
See whether AI assistants (ChatGPT, Claude, Perplexity, Google AI) actually mention a brand in their answers, and how that share-of-voice is trending versus competitors, week over week. Call with NO argument to list the user's monitored brands with each one's current mention rate and direction; pass a brand name or project id to get that brand's full trend, per-provider breakdown, competitor comparison, average position, and short example answers. Reads data the user's LLM Monitor projects have already collected — it does not trigger new runs. Read-only. Requires an API key.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Brand name or project id to detail. Omit to list all monitored brands. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint and openWorldHint, so the description correctly reinforces 'Read-only' rather than contradicting it. It adds valuable context beyond annotations: it reads previously collected LLM Monitor data, does not trigger new runs, and requires an API key. This exceeds the minimum bar for behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than the examples but every sentence carries functional weight: purpose, invocation modes, behavioral scope, and auth requirement. It is front-loaded with the core purpose and avoids filler, though it could be slightly tightened without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with one optional parameter and no output schema, the description covers the essential operational context: what data is read, what the two call patterns return, and the API key requirement. It does not mention pagination, error cases, or exact formatting of the project id, but these are minor for this simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema description coverage is 100%, the description meaningfully enriches the sole parameter. It explains the omission behavior (list all monitored brands) and the provided behavior (full trend, per-provider breakdown, competitor comparison, average position, example answers) — details far beyond the schema's one-line parameter description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a highly specific verb and resource: seeing whether AI assistants mention a brand and how that share-of-voice trends versus competitors. It also clearly distinguishes the two modes of operation (list all vs. detail one), making it easy to tell apart from the unrelated sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete instructions on when to call with no argument versus with a brand name/project id, and clarifies that it does not trigger new runs. It does not explicitly name alternatives or exclusion conditions, but the sibling tools are thematically distinct enough that no routing ambiguity arises.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_robotsValidate robots.txtARead-only
Audit a robots.txt file line-by-line. Detects syntax errors, empty user-agent groups, orphan Allow/Disallow lines, non-slash paths, non-numeric Crawl-delay values, unofficial Noindex usage, and wildcard traps. Also summarizes AI-bot coverage across 50+ known AI crawlers. Provide EITHER url (to fetch and audit) OR text.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | URL of the site whose /robots.txt should be fetched and audited. | |
| text | No | Raw robots.txt text to audit directly (alternative to url). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true and openWorldHint=true. The description adds behavioral context beyond that by explaining that a URL is fetched and audited, while raw text is audited directly, and by listing the specific issues detected. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The first sentence states the action and scope, the second lists detection categories and input requirements. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers input modes, audit scope, and specialized AI-bot coverage, which is enough for an agent to invoke the tool correctly. The only minor gap is the lack of an explicit output format, but no output schema exists and the description still conveys that results are audit findings and a coverage summary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds value by stating the either/or relationship between url and text, and by clarifying that url means 'fetch and audit' while text means raw input, which goes beyond the schema property descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Audit'), names the resource ('robots.txt'), and enumerates concrete checks such as syntax errors, empty user-agent groups, orphan Allow/Disallow lines, and AI-bot coverage. This clearly distinguishes it from siblings like validate_schema and check_ai_readiness.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: use this tool when you need to audit a robots.txt file, and it explains the two input modes ('Provide EITHER url ... OR text'). It does not explicitly name alternative tools or when not to use them, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
validate_schemaValidate JSON-LD SchemaARead-only
Validate all JSON-LD structured data on a URL. Extracts every block, runs each through a rules engine covering 20+ common types (Article, Organization, Product, LocalBusiness, FAQPage, Recipe, Event, etc.), and reports required-field errors, recommended-field warnings, and type-specific gotchas.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to validate structured data on. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds meaningful behavioral detail: it extracts every script block, runs each through a rules engine covering 20+ types, and reports required-field errors, warnings, and type-specific gotchas. This goes beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no redundancy. It front-loads the primary purpose, then details the scope, process, and output in a compact format. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only tool with no output schema, the description is complete. It explains what is validated, how (extraction and rules engine), and what is reported (errors, warnings, gotchas). An agent has enough to call it correctly without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only one parameter (url) and 100% schema description coverage, the schema already fully documents the parameter. The description's mention of 'on a URL' adds no new syntax or format details. Per the rubric, the baseline of 3 is appropriate since the schema carries the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (validate) and resource (all JSON-LD structured data on a URL), and lists the covered types (Article, Organization, Product, etc.), which clearly distinguishes it from sibling tools like generate_schema or validate_robots. It leaves no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool: whenever you need to validate JSON-LD structured data on a given URL. However, it does not explicitly mention alternatives or state when not to use it (e.g., for generating schema or validating robots.txt). The context is clear but lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.2.3- First observed
check_ai_readiness - First observed
check_content_parity - First observed
generate_llms_txt - First observed
generate_robots_txt - First observed
generate_schema - First observed
get_monitor_trend - First observed
validate_robots - First observed
validate_schema
TDQS
Scored across 8 tools
Each tool targets a distinct phase or concern: readiness scoring, schema validation, robots auditing, content parity, file generation, and trend monitoring. The only mild ambiguity is check_ai_readiness, which overlaps at a high level with the more specialized validators, but its first-step audit role is clearly described.
All tools follow a lowercase snake_case verb_noun pattern. The mix of check_, validate_, generate_, and get_ verbs is mostly predictable, though check and validate are close synonyms that create a minor stylistic inconsistency.
Eight tools is a well-scoped size for an AI crawl readiness server. Each tool covers a meaningful capability without redundancy, and the count supports both auditing and fixing workflows.
The set covers the main audit workflow (readiness, schema, robots, content parity), generation fixes (llms.txt, robots.txt, schema), and a monitoring view. Minor gaps exist—such as no validator for generated llms.txt and no detailed meta-tag inspection—but agents can complete core tasks without dead ends.
Maintenance
Related MCP Connectors
Run SEO + AI-visibility (GEO) audits from Claude, Cursor & other AI clients.
Free MCP tools for AI-search visibility: crawler checks, page audits, and llms.txt generation.
Audit any site's AI visibility from your assistant: crawler access, rendering, and schema.
Live SEO workflow tools for Claude Code, Codex, and AI agents.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceMCP server exposing Generative Engine Optimization tools to any AI agent, enabling checking of llms.txt, auditing robots.txt for AI crawlers, and validating JSON-LD schema.MIT
- AlicenseAqualityBmaintenanceEnables chatting with your site's SEO using real data, including crawling, on-page auditing, Google Search Console performance, Core Web Vitals, and competitor comparison from MCP clients like Claude or Cursor.143 npm1MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to crawl live websites, audit AEO readiness, generate Schema.org @graph JSON-LD, llms.txt, ai.txt, and robots.txt, inject structured data into HTML, validate optimizations, and retrieve framework-specific code snippets.MIT

agentbuiltofficial
AlicenseNot gradedqualityCmaintenanceEnables AI agents to run a free AI-readiness audit of any URL, checking AI crawler rules, JavaScript-free page text, JSON-LD, llms.txt, sitemap, meta description, FAQ schema, and returning a score with findings and fixes. Also exposes the same capability via an A2A agent endpoint.MIT