Skip to main content
Glama
competlab

competlab-mcp-server

by competlab

check_ai_crawlers

Read-only

Check live which AI assistants may fetch a domain's pages from robots.txt, with per-assistant verdicts and the crawlers that decided each.

Instructions

Live check of which AI ASSISTANTS can fetch a site's pages, read from its robots.txt. assistantAccess is the answer — one verdict per assistant (ChatGPT, Claude, Perplexity, Microsoft Copilot, Google AI Overviews, Gemini Apps) with the crawlers that decided each named beside it. Count that array for totals; no count is stored, and there is deliberately no overall score. TWO THINGS IT DOES NOT TELL YOU, both of which get misreported: it says an assistant is PERMITTED to fetch the site, never that it cites it; and modelTrainingAccess is a separate, NEUTRAL fact — blocking training crawlers costs no visibility and is a legitimate content decision, so never report it as a gap or advise undoing it. The one exception is mechanical: where a token under modelTrainingAccess[].decidedByCrawlers also appears under assistantAccess[].decidedByCrawlers (Google-Extended is the documented case), that block DOES cost visibility — match on userAgentToken before applying the general rule. crawlers[].ruleAudience tells you whether a rule NAMED the crawler or a User-agent: * catch-all swept it up; the second is usually accidental and is the more actionable finding. When robots.txt cannot be read, it returns the read outcome and NO verdict — no assistant access, no crawler list, no advice — so check robotsTxt.read first. A failed read is not an open site.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
domainYesDomain to scan, e.g. example.com
industryNoIndustry context for benchmark comparison.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv4.0.1
    • removedInput schema / additionalProperties
      Removed value: -false
  2. Changed2 schema fields changedv3.0.0
    • changedInput schema / properties / industry / description
      Previous value: -"Industry context for benchmark comparison. Current values: news-media, arts-entertainment, law-government, finance-healthcare, saas-tech, ecommerce, other. The backend may add new values over time; pass any of the listed strings (or a future one) and the API will validate."New value: +"Industry context for benchmark comparison."
    • addedInput schema / properties / industry / enum
      Added value: +[
      +  "news-media",
      +  "arts-entertainment",
      +  "law-government",
      +  "finance-healthcare",
      +  "saas-tech",
      +  "ecommerce",
      +  "other"
      +]
  3. Addedv1.2.0

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover readOnlyHint/openWorldHint, but the description layers on substantial behavioral detail: the unreadable-robots.txt failure mode returns NO verdict, 'permitted' is explicitly not 'cited', modelTrainingAccess is neutral and must not be reported as a gap, and there is a mechanical exception when a token appears in both decidedByCrawlers lists. This is exactly the kind of context annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first clause and each subsequent sentence carries a distinct anti-misreporting rule, so little is wasted. It is dense and long for a 2-param read tool, but the length is paying for interpretation that the missing output schema would otherwise have provided.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the full return-value burden and does so: it names assistantAccess, decidedByCrawlers, modelTrainingAccess, userAgentToken, ruleAudience, and robotsTxt.read, plus the empty-verdict case. An agent has everything needed to call it and interpret the response correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both parameters (domain, industry) are documented there, including the enum values for industry. The description adds nothing about either parameter (industry is never mentioned), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+mechanism: 'Live check of which AI ASSISTANTS can fetch a site's pages, read from its robots.txt.' That is a precise enough scope (robots.txt-derived AI assistant fetch access) to separate it from siblings like check_sitemap, fetch_url, or get_ai_visibility_dashboard without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The bulk of the description is interpretation guidance (how to read assistantAccess, modelTrainingAccess, ruleAudience) rather than when-to-use routing. No alternative tool is named and no exclusion condition or prerequisite is given, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.