Skip to main content
Glama

Cite Files

citefiles_crawler_access

Test how a website answers the real AI crawlers — GPTBot, ClaudeBot, PerplexityBot and Google-Extended — by requesting its homepage as a browser and again as each crawler, then comparing. Catches CDN and WAF rules that block assistants without the owner knowing. Use when a site looks correct but is not being cited.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesThe public website address.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does reasonably well: it discloses that the tool issues real requests to the homepage both as a browser and as each of four named crawlers, and that it compares results. It omits rate-limit, permission, or side-effect notes, but the read-only nature of the fetch is implied by the described behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences ordered as what it does, why it matters, and when to use it, with no filler. The mechanism and the value proposition are both front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter diagnostic tool with no output schema and no annotations, the description explains the operation and its purpose well. It stops short of describing what the comparison actually returns (e.g., per-crawler pass/blocked status), which would fully close the gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter with 100% schema description coverage, so the baseline is 3. The description only implies that 'url' points at the homepage being tested and adds no syntax, scheme, or scope detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (test/compare), the exact resource (a site's response to named AI crawlers), and the mechanism (request as browser vs. as each crawler). It is clearly distinguishable from siblings like citefiles_scan or citefiles_check_files, which do not involve per-crawler UA comparison.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit trigger condition: 'Use when a site looks correct but is not being cited.' That is concrete context, but it does not name an alternative tool or state when NOT to use it, so it stops short of full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources