Skip to main content
Glama

TunnelMind Data API

intel_robots

Retrieves the target domain's robots.txt file and parses it for AI crawler disallow rules. Specifically detects policies for known AI crawlers (GPTBot, ClaudeBot, CCBot, Bytespider, etc.) and returns a structured summary of the crawling policy.

Use this tool when:

  • You need to know whether a domain has opted out of AI training data collection.

  • You want to check if a specific AI crawler is blocked before citing the domain.

  • You are building a dataset of AI-accessible vs AI-blocked domains.

Do NOT use this tool when:

  • You want training opt-out signals beyond robots.txt (TDM reservation, noai meta) — use intel_optout instead.

  • You want the full technology stack — use intel_stack instead.

  • You need tracker database data — use get_domain instead.

Inputs:

  • domain (query, required): Domain to probe.

Returns:

  • robots_txt_found: false if the domain returned 404 or the file is empty.

  • ai_crawlers_blocked: list of AI crawler user-agent names that are disallowed.

  • all_blocked: true if User-agent: * with Disallow: / is present.

  • raw: first 4096 characters of the robots.txt file.

Cost:

  • Free. No API key required.

Latency:

  • Typical: 1-2s, p99: 6s.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
asyncNoWhen true, return a task handle immediately instead of blocking. Poll get_task for the result.
domainYes
receiptNoWhen true, attach a signed Receipt v1.0 committed to the transparency log. Additive — a signing failure never costs you the observation (ADR-014).

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral burden. It discloses cost ('Free. No API key required.'), latency (typical 1-2s, p99 6s), return value details, and the truncation of raw output to 4096 characters. It does not explicitly mention rate limits or other failure modes, but the disclosed information is meaningful and goes beyond what annotations would provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections: overview, use cases, exclusions, inputs, returns, cost, latency. Every section adds useful information without redundancy or padding. It is longer than a minimal description but earns its length by providing actionable detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema, the description enumerates the exact return fields (robots_txt_found, ai_crawlers_blocked, all_blocked, raw) with explanations, which is essential for an agent to interpret results. Combined with usage guidance, cost, latency, and parameter info, the description fully compensates for missing annotations and schema detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67% (domain lacks a description, async/receipt have descriptions). The description adds context for the primary `domain` parameter ('Domain to probe') but does not mention async or receipt at all. The schema already covers those, and the description's domain context is marginal, so the value added beyond structured fields is limited.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description opens with a specific verb+resource: 'Retrieves the target domain's robots.txt file and parses it for AI crawler disallow rules.' It clearly distinguishes from sibling tools by naming alternatives like intel_optout, intel_stack, and get_domain. This is precise and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'Use this tool when' and 'Do NOT use this tool when' sections give concrete decision criteria and name specific alternative tools (intel_optout, intel_stack, get_domain). This is exactly the level of guidance needed for an agent to choose correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

TDQS

B3.3/5.0
Disambiguation2/5

Many tools overlap in purpose, such as cross_lens_verify, cross_lens_lookup, profile_entity, and preflight_should_i_act, which all return node verdicts with subtle differences. Sigil verification tools and receipt-related tools also have similar names and require deep reading to distinguish.

Naming Consistency3/5

The tool names are mostly readable, but the pattern is mixed: some use verb_noun (get_domain, create_subscription) while others use domain prefixes (sigil_*, ghostroute_*, intel_*). Within each domain, naming is consistent, but the overall style lacks uniformity.

Tool Count1/5

With 90 tools, this server is extremely overloaded. Even for a multi-purpose data API, the sheer number overwhelms and makes navigation difficult, far exceeding the typical well-scoped MCP server. The count is an extreme mismatch for the apparent scope.

Completeness4/5

The tool surface is very comprehensive, covering tracker lookup, cross-lens verification, receipts, compliance, subscriptions, tasks, intel probes, and more. Minor gaps exist, such as no batch cross-lens verification, but core workflows are well covered.

Resources