Skip to main content
Glama

AI crawler access audit

ai_crawler_audit
Read-only

Audit which AI crawlers a site allows or blocks in robots.txt — sorted by CONSEQUENCE: blocking a VISIBILITY crawler (OAI-SearchBot, PerplexityBot, ChatGPT-User, Claude-User…) removes the site from live AI answers, while blocking a TRAINING crawler (GPTBot, CCBot, Google-Extended…) only opts out of model training. Those need opposite decisions. Checks 25 known AI crawlers + whether /llms.txt exists, plus general robots.txt hygiene (Sitemap directives, whole-site Disallow foot-guns). Free (no LLM calls).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
domainYesRoot domain to audit, e.g. 'yoursite.com'.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=true, which the description agrees with (an audit is read-only; openWorldHint fits the free/no-LLM-calls claim). Beyond annotations, the description adds valuable behavioral traits: 'Free (no LLM calls)' signals cost/performance characteristics, and the scope disclosure (25 crawlers, /llms.txt, hygiene checks) sets accurate expectations for what the audit covers. No contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first clause, and the structure flows logically: purpose → consequence framework → scope enumeration → cost signal. Every sentence earns its place. It runs slightly long due to the parenthetical crawler name lists and the em-dash framing, but both carry genuine decision-making value for the agent. Minor trimming possible without losing substance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with readOnly/openWorld annotations, this is nearly complete: it explains scope, the decision framework, and cost profile. The only gap is the absence of any description of the output/return format — and since there is no output schema, the description carries that burden entirely. This is a minor omission against an otherwise thorough definition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — the single 'domain' parameter is fully documented in the schema ('Root domain to audit, e.g. 'yoursite.com''). The description references the site implicitly but adds no parameter-level detail beyond the schema. Baseline 3 is appropriate since the schema already carries the full burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Audit which AI crawlers a site allows or blocks in robots.txt.' It further distinguishes itself by enumerating its exact scope (25 known AI crawlers, /llms.txt existence, robots.txt hygiene), which clearly separates it from the sibling ai_visibility_report and other SEO tools in the list. The purpose is unambiguous and non-tautological.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides strong implied context — it frames the audit around the visibility-vs-training crawler decision, making it clear this is a pre-decision check before blocking or allowing crawlers. However, it never explicitly states when to choose this tool over siblings like ai_visibility_report or site_audit, nor does it name alternatives or exclusions. Usage is inferable but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.