Cite Files
Server Details
Check whether AI assistants can reach, read and cite a public website. No account for 4 of 5 tools.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
TDQS
Scored across 5 tools
Each tool targets a distinct stage or artifact: scan (new audit), check_files (served discovery files), crawler_access (crawler-specific responses), fix_prompt (remediation plan), and agent_commerce (stored commerce-readiness audit). The scan vs. agent_commerce overlap is clarified by descriptions, since one runs an audit and the other only reads a stored report.
All names share the citefiles_ prefix, but suffixes mix verb_noun (check_files, fix_prompt), noun phrases (agent_commerce, crawler_access), and a bare term (scan). This is readable and mostly predictable, though not a strict verb_noun pattern throughout.
Five tools is well-scoped for an AI-readiness auditing server. Each covers a distinct phase—scan, file verification, crawler access, remediation planning, and commerce audit—with no obvious filler.
The set covers audit, diagnosis, served files, crawler blocking, remediation planning, and commerce readiness. It lacks a direct tool to apply or generate fixes, so agents must rely on fix_prompt and human/TODO instructions, but this is a minor gap rather than a dead end.
Available Tools
5 toolscitefiles_agent_commerceAInspect
Requires an API token (Authorization: Bearer cf_…). Read the stored Agent Commerce Readiness audit for a scan on your citefiles account: whether an AI agent can authenticate without a human's password, discover an API, read prices as text, and be offered a machine-payable challenge — as a tier from T0 to T5. Create a token at https://citefiles.com/account. Only returns audits for scans your account owns. This tool never runs a new audit — run it from the scan's report page first, then call this tool to read the result.
| Name | Required | Description | Default |
|---|---|---|---|
| scan | Yes | The scan's public id, from a citefiles.com/report/<id> URL. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it declares the auth requirement and format (Authorization: Bearer cf_…), where to obtain a token, the ownership scoping rule, and the read-only/non-mutating behavior ('never runs a new audit'). It also sketches the returned judgment scale (tier T0 to T5), which is exactly the behavioral context missing from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, all of which carry information: auth, purpose, ownership scope, and the no-new-audit constraint. The only structural nit is that the opening sentence is an auth preamble, so the actual purpose arrives in the second sentence rather than being front-loaded, which slightly weakens scanability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with no annotations and no output schema, the description supplies everything an agent needs: credentials, prerequisite ordering, ownership limitation, and even the shape of the result (a T0–T5 tier) so the return value needn't be documented separately. Nothing material is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there is only one parameter, so the schema already explains that 'scan' is the public id from a citefiles.com/report/<id> URL. The description corroborates the source of that id ('run it from the scan's report page first') but adds no syntax or format detail beyond the schema, making the baseline 3 correct here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Read the stored Agent Commerce Readiness audit for a scan on your citefiles account', and it scopes the content of that audit (agent auth, API discovery, text-readable prices, machine-payable challenge, tier T0–T5). It also explicitly separates itself from the scan-running sibling by stating 'This tool never runs a new audit', so an agent can tell it apart from citefiles_scan without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear precondition and ordering: run the audit from the scan's report page first, then call this tool, and it notes the account-ownership restriction ('Only returns audits for scans your account owns'). What it lacks is explicit routing to an alternative when a scan has no audit yet — it implies raising the audit first but never names the tool that does it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
citefiles_check_filesAInspect
Check which AI-facing discovery files a website actually serves: llms.txt, llms-full.txt, ai.txt, robots.txt, sitemap.xml, AGENTS.md, .well-known/mcp.json and others. Distinguishes a real file from a single-page app returning its HTML shell for every unknown path, which is the usual reason a file appears to exist but does not. Use this to verify a file you just published is being served correctly.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The public website address. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it delivers a genuinely useful trait: it distinguishes a real file from an SPA returning its HTML shell for unknown paths. That is exactly the kind of non-obvious behavior an agent needs. It stops short of describing the response shape or whether checks are batched per-path, so it isn't exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action, then the differentiating behavioral insight, then the usage cue. No filler, and the most decision-relevant material comes first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, no-annotation, read-style check with no output schema, the description covers purpose, the key false-positive pitfall, and a use case. It would be fully complete if it hinted at what the result reports (which files present vs absent), but nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema description coverage, so the schema already documents 'url' fully. The description adds no syntax, format, or scope detail (e.g. whether paths or a bare origin are accepted) beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Check which ... files a website actually serves') and enumerates the resource set (llms.txt, robots.txt, sitemap.xml, MCP manifests). It cleanly separates itself from a generic scan by focusing on served discovery files, but it never names or contrasts against the sibling tools (citefiles_scan, citefiles_crawler_access), so sibling differentiation is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear triggering context: 'Use this to verify a file you just published is being served correctly.' That tells the agent when to reach for this tool, but there is no guidance on when not to use it or which sibling to prefer for the adjacent jobs (scanning, crawler access), so it falls short of the explicit alternatives level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
citefiles_crawler_accessAInspect
Test how a website answers the real AI crawlers — GPTBot, ClaudeBot, PerplexityBot and Google-Extended — by requesting its homepage as a browser and again as each crawler, then comparing. Catches CDN and WAF rules that block assistants without the owner knowing. Use when a site looks correct but is not being cited.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The public website address. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does reasonably well: it discloses that the tool issues real requests to the homepage both as a browser and as each of four named crawlers, and that it compares results. It omits rate-limit, permission, or side-effect notes, but the read-only nature of the fetch is implied by the described behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences ordered as what it does, why it matters, and when to use it, with no filler. The mechanism and the value proposition are both front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter diagnostic tool with no output schema and no annotations, the description explains the operation and its purpose well. It stops short of describing what the comparison actually returns (e.g., per-crawler pass/blocked status), which would fully close the gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema description coverage, so the baseline is 3. The description only implies that 'url' points at the homepage being tested and adds no syntax, scheme, or scope detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (test/compare), the exact resource (a site's response to named AI crawlers), and the mechanism (request as browser vs. as each crawler). It is clearly distinguishable from siblings like citefiles_scan or citefiles_check_files, which do not involve per-crawler UA comparison.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger condition: 'Use when a site looks correct but is not being cited.' That is concrete context, but it does not name an alternative tool or state when NOT to use it, so it stops short of full when/when-not routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
citefiles_fix_promptAInspect
Get a task list for making a website readable and citable by AI assistants, derived from a real scan of it. Returns the findings with the evidence behind each, and the rules to follow while fixing them — most importantly that you must not invent facts about the business, and must leave a visible TODO instead. Use this to plan the work after citefiles_scan.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The public website address. | |
| style | No | Prose instruction (claude) or a task list with acceptance criteria (codex). Defaults to claude. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does well: it discloses that output is derived from a real scan, includes findings plus supporting evidence, and embeds operating rules — notably not inventing business facts and leaving a visible TODO. It does not cover permissions, rate limits, or whether the call is purely read-only, but the behavioral content is richer than most.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core purpose and ending with the routing instruction. The parenthetical warning about inventing facts is somewhat dense but earns its place by conveying a key constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description appropriately summarizes the return (findings with evidence and fixing rules). For a two-param read tool with no annotations, this is nearly complete; it lacks only explicit read-only confirmation and handling of the style parameter's effect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (url, style enum with default) are already documented in the schema. The description adds no meaning about parameter usage, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (get) and resource (a task list for making a website readable/citable), plus the origin of the content ('derived from a real scan'). It distinguishes itself from the sibling citefiles_scan by positioning itself as the downstream planning step, though it doesn't explicitly contrast with the other siblings (check_files, crawler_access, agent_commerce).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this to plan the work after citefiles_scan' gives explicit sequencing relative to a named sibling, which is clear usage context. There is no explicit when-not guidance or coverage of the other siblings, keeping it short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
citefiles_scanAInspect
Check whether AI assistants can reach, read and cite a public website. Returns a citation-readiness score out of 100, the category breakdown, and the specific problems found — crawler blocks, missing llms.txt/ai.txt/sitemap, thin or JavaScript-only content, missing structured data. Use this before changing a site for AI visibility, and again afterwards to check the change worked. Scans of the same site within six hours are reused.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The public website address, e.g. example.com or https://example.com |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose a non-obvious behavior: scans of the same site within six hours are reused, which matters for interpreting staleness. It also enumerates what is reported (score, category breakdown, specific problems). It stops short of stating rate limits or any authentication expectations, though the tool targets public sites.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose and output first, then usage timing, then the caching caveat. Front-loaded and free of filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description adequately describes return values (score out of 100, category breakdown, specific problems) and covers the caching behavior. For a single-parameter public read tool, little else is needed, though a note on scan duration or freshness semantics beyond the six-hour reuse would round it out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single url parameter is fully documented in the schema (100% coverage), including format examples, so the baseline is 3. The description only adds that the site must be public, which is a marginal gain beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Check whether AI assistants can reach, read and cite a public website') and states the concrete output: a citation-readiness score out of 100 plus a category breakdown. It does not explicitly distinguish itself from siblings like citefiles_crawler_access or citefiles_check_files, whose concerns (crawler blocks, missing llms.txt/sitemap) it also mentions, so an agent must infer that this is the aggregate scan.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit timing guidance: use before changing a site for AI visibility and again afterwards to verify the change. However, it names no alternatives and offers no exclusions, so an agent is not told when to prefer the narrower sibling tools over this broad scan.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
- First observed
citefiles_agent_commerce - First observed
citefiles_check_files - First observed
citefiles_crawler_access - First observed
citefiles_fix_prompt - First observed
citefiles_scan
Related MCP Connectors
Check if ChatGPT, Claude and Perplexity can reach, read and quote a website.
Can ChatGPT, Claude and Perplexity reach and read a website? Score, problems found and fixes.
Checks whether a website is readable and citable by AI systems (ChatGPT, Claude, Perplexity, etc.)
Check how well AI agents can discover and understand a public website.
Related MCP Servers
AlicenseNot gradedqualityCmaintenanceChecks whether a website is readable and citable by AI search engines — llms.txt, Schema.org structured data, AI-bot access in robots.txt, content freshness, answer directness, E-E-A-T signals, plus a LocalBusiness Rich Results validator. Free, no API key, remote Streamable HTTP.1MIT- AlicenseAqualityCmaintenanceCheck whether a website is visible to AI search engines (ChatGPT, Perplexity, Claude, Google AI Overviews). Returns a 0-100 readiness score, a grade, and a specific fix for each gap. Dependency-free, no API keys.23 npmMIT

citedbyai-mcp-serverofficial
AlicenseNot gradedqualityBmaintenanceFree AI citation readiness checker powered by Cited By AI's CPS® framework. Instantly scores any website 0-100 across structured data, meta tags, content quality, technical config, and AI signals. Returns a grade (A-F) and the top issues blocking AI citation in ChatGPT, Claude, Perplexity, Gemini, and Copilot. No auth required.1MIT- AlicenseNot gradedqualityCmaintenanceLets an AI assistant check any URL to see how much of a page's main content is present in the raw HTML versus only after JavaScript runs, showing which text blocks, meta fields, canonical tags, JSON-LD and links AI crawlers like GPTBot and ClaudeBot cannot read. It also reports what robots.txt allows for 22 AI bot tokens and how the site responds to each bot's user-agent.AGPL 3.0
Glama MCP Gateway
Add one secure layer between your agents and this server.