Siteiz
Server Details
Check which AI crawlers a site allows and see pages the way AI crawlers do. Free, read-only.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
TDQS
Scored across 4 tools
Each tool targets a distinct aspect: check_ai_crawlers covers robots.txt permissions, view_page_as_ai_crawler covers actual rendered content, explain_ai_crawler is educational, and get_ai_visibility_report retrieves report data. The two crawler-checking tools share the same 'AI crawler' domain but their descriptions clearly separate permission-checking from content rendering.
All four tools use lowercase snake_case with a consistent verb_noun structure (check_, explain_, get_, view_). The pattern is predictable and verbs are distinct actions.
Four tools is a tight, focused set for an AI-visibility diagnostic niche, with each tool earning its place. It is slightly thin—there is little room for deeper analysis—but not problematically under-scoped.
The surface covers robots.txt permissions, crawler education, content rendering, and published reports, but get_ai_visibility_report only works for well-known companies, leaving a notable gap for auditing arbitrary sites. No tool lets an agent generate or act on remediation (e.g. llms.txt creation or visibility scoring for arbitrary URLs).
Available Tools
4 toolscheck_ai_crawlersCheck AI crawler accessARead-onlyIdempotentInspect
Reads a website's robots.txt and reports, for 11 AI crawlers (GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, PerplexityBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider), whether each may reach the given page. Separates AI search crawlers (decide if the site appears in ChatGPT, Claude and Perplexity answers) from training crawlers. Also reports llms.txt and declared sitemaps.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Website or page address, e.g. example.com or https://example.com/pricing |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | Yes | |
| agents | Yes | |
| llmsTxt | No | |
| sitemaps | No | |
| robotsTxt | Yes | |
| httpStatus | No | |
| contentSignals | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, covering the safety and idempotency profile. The description adds useful scoping context (reads robots.txt, reports 11 named crawlers, llms.txt and sitemaps) for an open-world network read, but doesn't mention rate limits, caching, redirects, or failure modes when robots.txt is missing. With annotations carrying the safety burden, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with zero filler, front-loaded on the core action and then the output taxonomy (search vs training crawlers, llms.txt, sitemaps). Every sentence earns its place and the crawler list is information the agent needs, not padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return shape needn't be explained, and the description still previews the key result categories (per-crawler verdicts, search vs training split, llms.txt, sitemaps). It is complete enough to call correctly; missing only edge-case behavior (no robots.txt, timeouts) which is minor for a read-only checker.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Single parameter with 100% schema coverage, so the schema already documents the url format with an example. The description adds the intent that the url points at a 'given page' whose reachability is being tested, which clarifies semantics slightly beyond the schema. Baseline for a well-covered single param, nudged up by that framing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource: reads robots.txt and reports per-crawler access for a named page. It names the exact 11 crawlers and the llms.txt/sitemap outputs, so an agent knows exactly what comes back without opening the output schema. It is distinguishable from siblings like explain_ai_crawler and view_page_as_ai_crawler by being the raw robots.txt access check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case (determining whether AI crawlers may reach a page) and even hints at the search-vs-training distinction, but it never explicitly says when to choose this tool over explain_ai_crawler or view_page_as_ai_crawler. The routing between these four siblings is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
explain_ai_crawlerExplain an AI crawlerARead-onlyIdempotentInspect
Explains what an AI crawler user agent does (operator, purpose, and what blocking it in robots.txt changes). Pass a name such as GPTBot or OAI-SearchBot, or omit it to list all crawlers Siteiz tracks.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Crawler user agent, e.g. GPTBot. Omit to list all. |
Output Schema
| Name | Required | Description |
|---|---|---|
| crawlers | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered. The description adds only that the tool explains robots.txt blocking consequences and that omitting the name enumerates tracked crawlers — useful but modest, and the enumeration behavior is already stated in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The core purpose and its subfields come first, and the calling instruction follows — well front-loaded and tightly sized for a one-parameter lookup tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a rich annotation set, an output schema, and one optional parameter, the description covers what the tool explains and how to call it. The only gap is guidance on choosing this tool over its siblings, which is minor for a simple explanatory lookup.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single optional parameter with 100% schema description coverage, so the baseline is 3. The description supplies concrete examples (GPTBot, OAI-SearchBot) and the omit-to-list-all semantics, but the schema already documents both.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Explains') and resource ('AI crawler user agent') and enumerates the payload: operator, purpose, and the effect of robots.txt blocking. That content profile clearly separates it from check_ai_crawlers, get_ai_visibility_report, and view_page_as_ai_crawler, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Pass a name such as GPTBot or OAI-SearchBot, or omit it to list all crawlers Siteiz tracks' gives invocation guidance for the parameter, but says nothing about when to reach for this tool over check_ai_crawlers or the visibility report. Usage is implied rather than framed against alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_ai_visibility_reportGet a published AI visibility reportARead-onlyIdempotentInspect
Returns a published, dated Siteiz AI visibility report for a well-known company's homepage (score, grade, pillar scores, top issues). Omit the company to list all published reports.
| Name | Required | Description | Default |
|---|---|---|---|
| company | No | Company name, e.g. Notion. Omit to list all reports. |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | No | |
| grade | No | |
| score | No | |
| report | No | |
| company | No | |
| pillars | No | |
| reports | No | |
| scannedAt | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover readOnlyHint, openWorldHint, idempotentHint, destructiveHint. The description adds that only published reports are returned and that omitting company lists all reports. With annotations already covering safety, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the primary return and the optional list mode. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given one optional parameter and an output schema, the description is nearly complete. It could mention report date range or pagination for the list mode, but is otherwise sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds that omission triggers a list action and that reports are for well-known companies, adding meaning beyond the schema. Thus a 4 is warranted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: returns a published, dated AI visibility report for a company homepage. Distinguishes from siblings (crawler tools) by naming the report contents (score, grade, pillar scores, top issues).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states the dual mode: omit company to list all reports. Provides clear context. No explicit when-to-use vs when-not, but the dual behavior is clearly explained.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
view_page_as_ai_crawlerView a page as an AI crawlerARead-onlyIdempotentInspect
Fetches one page the way most AI crawlers do (plain HTTP, no JavaScript) and reports what they actually receive: title, meta description, canonical, H1, heading count, structured data types, visible word count, token reading cost, and a text preview. Use it to check whether content is visible to AI without JavaScript.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Website or page address, e.g. example.com or https://example.com/pricing |
Output Schema
| Name | Required | Description |
|---|---|---|
| h1 | No | |
| bytes | No | |
| title | No | |
| words | Yes | |
| status | Yes | |
| finalUrl | Yes | |
| fullScan | Yes | |
| canonical | No | |
| htmlTokens | No | |
| textTokens | No | |
| headingCount | No | |
| structuredData | No | |
| metaDescription | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/openWorld, so the safety profile is covered. The description adds genuinely non-annotated behavior: the fetch emulates crawler-style plain HTTP with no JavaScript execution, which tells the agent the result reflects a JS-less rendering rather than a real browser. It stops short of covering errors, redirects, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the mechanism and use case front-loaded before the field list. The response field enumeration is somewhat redundant against the output schema but is compact enough not to bloat the definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a 1-parameter schema, full annotations, and an output schema present, the description need not explain return values, yet it helpfully summarizes the payload shape. The remaining gap is sibling disambiguation, which matters given three adjacent visibility tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already gives format examples (example.com or https://example.com/pricing), so the single parameter is fully documented structurally. The description adds no additional constraint or format meaning beyond that, matching the baseline for high-coverage schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (fetches one page) plus the exact mechanism (plain HTTP, no JavaScript), which is more precise than the title alone. However, it never distinguishes itself from siblings like get_ai_visibility_report or check_ai_crawlers, which plausibly cover adjacent ground.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use it to check whether content is visible to AI without JavaScript" gives one concrete scenario, so usage is implied rather than left blank. But with three sibling tools in the same visibility domain, the description offers no when-not condition and no routing guidance to an alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
4 tool updates
- First observed
check_ai_crawlers - First observed
explain_ai_crawler - First observed
get_ai_visibility_report - First observed
view_page_as_ai_crawler
Related MCP Connectors
Audit webpage access for AI crawlers. API key and Premium or Partner plan required.
Scan any website's AI readiness: AI search visibility and AI agent usability. Free, no auth.
Checks whether a website is readable and citable by AI systems (ChatGPT, Claude, Perplexity, etc.)
Free AI-readiness audit of any URL: AI crawler rules, JS-free text, JSON-LD, llms.txt. Tool catalog.
Related MCP Servers
- AlicenseAqualityCmaintenanceAnalyze and generate robots.txt files with AI crawler awareness. Fetch any site's robots.txt, detect which AI bots (GPTBot, ClaudeBot, PerplexityBot, Google-Extended) are blocked or allowed, and generate optimized robots.txt with toggle controls for 20+ AI crawlers.51MIT
- AlicenseAqualityDmaintenanceAudits AI-bot visibility: robots.txt per-bot for 22 AI user-agents (GPTBot/ClaudeBot/PerplexityBot/etc), Cloudflare flags, JSON-LD, sitemap, llms.txt, SPA shell, plus cross-model brand mentions via Perplexity + OpenRouter. 0-100 score. SSRF-guarded, spend-capped.41MIT
- AlicenseAqualityCmaintenanceEnables inspection of any website's AI-search readiness, checking AI crawler blocks, llms.txt, schema markup, and indexing directives from MCP clients like Claude.434 npmMIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to perform instant SEO audits, check robots.txt, sitemaps, and AI crawler access for any URL without API keys.MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.