Skip to main content
Glama

CrawlCheck

Server Details

Verification layer for the agentic web: a signed answer about any site before an agent acts.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
emmanuelorta/crawlcheck-core
GitHub Stars
0

TDQS

Score is being calculated.

Available Tools

23 tools
corpus_state
Read-onlyIdempotent
Inspect

Coverage of CrawlCheck's public dataset: how many distinct domains have been measured and how they split by platform, rendering, size, language and kind, with the strata that are still under-sampled.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

crawler_path
Read-onlyIdempotent
Inspect

Where each crawler's path into a domain dies, as an observed data flow: edge decision, robots.txt as a file, robots.txt rules, then the page, JSON-LD, sitemap, llms.txt and entitymap.json it reached. Give agent (e.g. ClaudeBot, GPTBot, googlebot) for one identity's path and a one-line answer; omit it for every identity plus the stores and breaks. Uses the latest record on file, or scans first when there is none (fresh=true forces a scan). The answer-engine output is drawn but never measured.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoIdentity id or label, e.g. claudebot, GPTBot, Googlebot
freshNoScan now instead of using the record on file
domainYes
draft_llms
Read-onlyIdempotent
Inspect

Draft an llms.txt from the domain's own homepage and up to 25 declared pages, using their titles and descriptions. A draft for the owner to cut down, not a publication.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYes
explain_finding
Read-onlyIdempotent
Inspect

Why CrawlCheck decided a finding: the rule and its revision, the exact fetches it compared (identity, status, bytes, sha256), when the condition held on the site, and a Merkle proof that the decision is inside the signed manifest. Give id (a d1: decision or f1: finding id from a report) or domain for its top finding. Other findings, and lookups by code, need a licence for that domain (send it in x-crawlcheck-key).

ParametersJSON Schema
NameRequiredDescriptionDefault
idNod1: decision id, f1: finding id, or report id
codeNoFinding code (licence only)
pathNo
domainNoBare domain; alone it explains the top finding
fix_entitymap
Read-onlyIdempotent
Inspect

A starter entitymap.json built from the domain's latest scan record: name, phone, coordinates and declared service areas as entities with SERVES relations. Fields the page never stated are marked TODO, never guessed. Needs a prior scan.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYes
fix_machine_chain
Read-onlyIdempotent
Inspect

A fix plan for MACHINE_CHAIN_HEAVY: the files a crawler reads before the page and the 404s absent agent files return, then measured reductions (llms.txt to an index, a sitemap index, robots.txt comments and duplicate groups, minified JSON, short 404s) with the chain size after each and the step at which the detector stops firing.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoReport id
domainNoUses the latest scan on file
fix_mostly_code
Read-onlyIdempotent
Inspect

A fix plan for PAGE_IS_MOSTLY_CODE: the homepage's measured bytes by category, the largest blocks with what to do with each, ordered steps with the text ratio after each, and the step at which the detector would stop firing - or, when moving code cannot clear it, how much markup to cut or text to add. Give domain (fetched now) or id (a stored scan).

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoReport id: plan from the stored scan instead of fetching
domainNo
fix_robots
Read-onlyIdempotent
Inspect

The domain's served robots.txt, corrected: * group rules copied into named groups that lacked them (shadowing), a Sitemap line added when missing, a minimal replacement when the served file was HTML. Every change is listed at the top; nothing else is touched.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYes
fix_stale_cache
Read-onlyIdempotent
Inspect

A fix plan for STALE_CACHE_SERVED: the homepage fetched now, its Age, the cache layers that name themselves in the headers in purge order (innermost first), the HTML lifetime each response declares, a fresh-read comparison, and whether a purge clears the finding and keeps it cleared.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoReport id; the page is still fetched now, because cache state is live
domainNo
list_verified_capabilities
Read-onlyIdempotent
Inspect

What a domain declares agents can do (OpenAPI operations, MCP tools, A2A skills, API catalog entries), each classed by safety (read_only_public up to financial), and which ones CrawlCheck actually called and confirmed - plus every mismatch between declaration and behaviour. Only read-only public actions are ever called; everything else says why it was not. From the latest scan on file.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYes
verified_onlyNoReturn only the capabilities that were confirmed
machine_record
Read-onlyIdempotent
Inspect

The unified machine record for one observation: subject, grade, findings (the top one in full without a licence), capabilities with verification, and the evidence roots, manifest and verifier links that let anyone check it offline. Give id (report, d1:, f1: or manifest sha256) or domain for the latest.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNo
domainNo
mcp_servers
Read-onlyIdempotent
Inspect

Which remote MCP servers the official MCP Registry lists on a domain, and what each endpoint actually answered when CrawlCheck sent the MCP handshake (initialized, auth_required, unreachable, tool count, tool-list drift). Use before selecting an MCP tool from that domain.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesBare domain or host, e.g. example.com
preflight
Read-onlyIdempotent
Inspect

Call BEFORE reading, citing, connecting to or transacting with a domain. Returns a signed decision: allow, warn, require_confirmation, block or unsupported, with every step that led there and the resolve answer it was made from. Pass your crawler token as agent so robots.txt is checked for you. A policy can only make the decision stricter. Do not proceed on block; ask the user on require_confirmation.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoYour crawler user-agent token, e.g. GPTBot
actionNoWhat you are about to do (default read)
domainYesBare domain, e.g. example.com
policyNomax_age_hours, on_warn (warn|require_confirmation|block), require_entity, human_approval[], allow[], deny[]
templateNosafe_citation, safe_data_retrieval, safe_api_connect, safe_mcp_tool, safe_commerce, safe_oauth or support_escalation
public_counts
Read-onlyIdempotent
Inspect

The figures CrawlCheck quotes about itself, read from the source: sites_measured, domains_in_corpus, crawler_visits (a rolling window), sections, sections_scored, findings_published, guides_published, with an at timestamp.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

registry_lookup
Read-onlyIdempotent
Inspect

Whether a domain is in CrawlCheck's registry and what it serves: llms.txt, agents.md, a media kit at /.well-known/media-kit.json, a reciprocity-tested entity graph, an AI access policy that names crawlers, and agent-callable surfaces. With no domain, returns the shelf counts across every measured site. A row exists because a fetch produced it; nothing here is self-reported and no payment moves a shelf.

ParametersJSON Schema
NameRequiredDescriptionDefault
shelfNoFilter the list to one shelf: llms, agents, mediakit, entity, aipolicy, agentapi
domainNoA bare domain. Omit for the corpus-wide shelf counts.
resolve_domain
Read-onlyIdempotent
Inspect

Call before fetching from, citing or acting on a domain. Returns one signed answer: crawler policy as declared (robots.txt per crawler), what each crawler identity was actually served, machine files, declared capabilities with their safety/policy class and whether any was verified, entity reciprocity, and open findings. Each section is declared, observed or not_measured with its time; there is no overall score. An unknown domain returns not_measured (never a guess) and is queued for measurement.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesBare domain, e.g. example.com
resolve_robots
Read-onlyIdempotent
Inspect

Resolve every named answer engine, search index and training crawler against a domain's robots.txt the way a crawler does: most-specific group only, longest match, allow wins a tie. Shows when a Disallow under * does not apply to an agent with its own group.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoPath to test, default /
domainYes
scan_domainInspect

Scan a public domain as several crawler identities and return the graded record: grade, AI-visibility score, per-section scores, findings with evidence, and whether the origin refused CrawlCheck. One scan takes 5-15 seconds.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYesA bare domain such as example.com, or a full URL
telemetry
Read-onlyIdempotent
Inspect

Verified-crawler traffic observed at crawlcheck.io itself: which named agents arrived, how many claims were confirmed against operator ranges, how many were forged.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

verify_crawler_log
Read-onlyIdempotent
Inspect

Given raw web-server access-log lines, decide for each line that names a crawler (GPTBot, ClaudeBot, Googlebot, PerplexityBot...) whether the source IP falls inside that operator's published ranges. Returns verified, spoofed, or unverifiable when the operator publishes no ranges.

ParametersJSON Schema
NameRequiredDescriptionDefault
linesYesUp to 200 raw log lines, newline separated
video_channel
Read-onlyIdempotent
Inspect

Every upload on a YouTube channel, read 50 at a time, sorted by views. The upload list is free; the per-check tally needs a Watch licence and is not offered here.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoChannel id, if no handle
handleNo@channel handle
video_check
Read-onlyIdempotent
Inspect

One YouTube video against the video anchor model: metadata, chapters, captions and storyboard where observed, twelve checks each naming what it reads.

ParametersJSON Schema
NameRequiredDescriptionDefault
vYesVideo id or URL
freshNoBypass the six-hour cache
video_site
Read-onlyIdempotent
Inspect

The site half of video: homepage plus up to 60 sitemap pages read for YouTube embeds, facades and page-builder widgets, VideoObject nodes and their required properties, and - with a handle - how many of the channel's videos the site carries.

ParametersJSON Schema
NameRequiredDescriptionDefault
domainYes
handleNo

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 23 tool updates
    • First observedcorpus_state
    • First observedcrawler_path
    • First observeddraft_llms
    • First observedexplain_finding
    • First observedfix_entitymap
    • First observedfix_machine_chain
    • First observedfix_mostly_code
    • First observedfix_robots
    • First observedfix_stale_cache
    • First observedlist_verified_capabilities
    • First observedmachine_record
    • First observedmcp_servers
    • First observedpreflight
    • First observedpublic_counts
    • First observedregistry_lookup
    • First observedresolve_domain
    • First observedresolve_robots
    • First observedscan_domain
    • First observedtelemetry
    • First observedverify_crawler_log
    • First observedvideo_channel
    • First observedvideo_check
    • First observedvideo_site

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Enables production-grade answer verification for LLM agents by independently re-checking answers with a configurable verifier model, enforcing confidence policies, and producing Ed25519-signed, auditable verification results with optional web search and knowledge-base evidence.
    3
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI agents to obtain independent, Ed25519-signed receipts confirming real-world facts: whether a page is online, a domain is legitimate, an email address is deliverable, a heartbeat was declared, or a scheduled task actually ran. Each response includes a verifiable cryptographic receipt, so agents can prove to their users that an asserted outcome was observed by a neutral third party rather than self-attested.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to check trustworthiness before recommending URLs, products, or organizations, with fail-closed pass/fail verdicts and attested-only recommendations.
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Provides AI agents with independent, advisory reviews before irreversible actions, returning recomputable, Bitcoin-anchored signed proofs that anyone can verify for free via a public verdict ledger.
    -
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.