CrawlBit MCP
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@CrawlBit MCPCan AI crawlers read my site example.com?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
CrawlBit MCP
Four free tools that answer one question: can AI engines find, read and recognise a website?
No account. No API key. No signup. Nothing here runs a language model, so nothing here costs you or us anything.
"Can ChatGPT read stripe.com?"
"Why does AI never mention my store?"
"Where is my brand missing outside my own site?"Ask in plain language. Your AI client picks the right tool.
Install
Requires Node.js 18+.
Claude Code
claude mcp add crawlbit -- npx -y crawlbit-mcpClaude Desktop
Add to claude_desktop_config.json:
{
"mcpServers": {
"crawlbit": {
"command": "npx",
"args": ["-y", "crawlbit-mcp"]
}
}
}Cursor
Add to .cursor/mcp.json:
{
"mcpServers": {
"crawlbit": {
"command": "npx",
"args": ["-y", "crawlbit-mcp"]
}
}
}Restart your client. Four tools appear.
Related MCP server: seo-audit-mcp
The four tools
Tool | What it answers |
| Can GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended read this site? |
| Does this brand read as a clear entity AI can recognise? |
| Where is this brand missing outside its own website? |
| Can AI shopping agents understand and recommend these products? |
Try these
Check whether AI crawlers can read example.com
Does example.com read as a clear brand entity to AI?
Where is example.com missing off-page, and which gap should we close first?
Compare example.com and competitor.com on AI crawler accessThe tools compose. A useful sequence is: crawler access first (a blocked crawler makes everything else pointless), then the entity check, then the off-page gaps.
What this does NOT do
Written plainly, because an SEO tool that overstates itself is worse than no tool.
It does not test whether AI actually cites a brand. That means running real buyer questions against ChatGPT, Perplexity and Claude, reading the answers, and reporting who gets named instead of you. It costs real money per run and is not part of this free surface. These four tools measure the conditions that make a citation possible, which is a smaller claim.
crawler_watchreadsrobots.txtonly. It cannot see blocking done at the CDN, firewall or rate-limit layer. A site can pass here and still refuse the crawler in practice.entity_checkreads what the site publishes about itself. It says nothing about the brand's reputation across the rest of the web.shoppinginspects published structured data. It cannot see a merchant feed submitted privately to a platform.
If a site scores well on all four and still is not cited, the answer is almost always the same and it is not technical: too few third-party pages mention it. These tools will tell you that honestly rather than sell you a fix that does not exist.
Why only four tools
CrawlBit also runs a full technical audit and generates llms.txt. Both use a language model,
so both cost money per run, and neither is exposed here.
That is a deliberate line rather than a teaser. A free tool whose bill grows with its popularity
gets rate-limited, degraded or withdrawn the moment it succeeds, and the people who installed it
are the ones who pay for that. Everything in this server costs nothing to run, so nothing here
has to be taken back later. The audit and the llms.txt generator live at
crawlbit.app if you want them.
Rate limit
The API rate-limits per IP. Normal conversational use does not come close, and the server runs on your machine, so the budget is yours rather than shared. If you do hit it, you get a plain sentence rather than a stack trace.
Configuration
Variable | Default | Purpose |
|
| Point the server at another instance |
Privacy
The server runs locally and stores nothing. The only data leaving your machine is the URL you ask about, sent to the CrawlBit API to be analysed. There is no account, so there is nothing to tie a request to a person.
About
Built by CrawlBit, which measures AI visibility and does the off-page work that earns citations. The free tools here are the measurement half. The paid work is the other half, and you are not required to look at it to use this.
MIT licensed. Issues and pull requests welcome.
Available Tools
4 toolscrawlbit_crawler_watchAI crawler accessA
Checks whether AI crawlers are allowed to read a site: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot and Google-Extended. Reads the site's robots.txt and reports, per crawler, whether it is fully allowed, partially blocked or blocked entirely. Use this when someone asks why AI never mentions their site, or before any other AI-visibility work: a blocked crawler makes everything else pointless. Reflects robots.txt only, so it cannot detect blocking done at the CDN or firewall layer.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The site to analyse, e.g. example.com or https://example.com |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well by disclosing that it 'Reads the site's robots.txt' and reports per-crawler statuses, while also stating the limitation that it only reflects robots.txt and not CDN/firewall blocking. This gives the agent a clear picture of what the tool does and cannot do, though details like response format or rate limits are omitted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences long and front-loaded with the core function, followed by usage guidance and a limitation. Every sentence contributes meaningful information with no redundancy or filler, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (single parameter, no output schema, no annotations), the description is reasonably complete. It covers what the tool does, when to use it, and a key limitation. While it does not explicitly detail the return format, the phrase 'reports, per crawler, whether it is fully allowed, partially blocked or blocked entirely' gives the agent enough context. A small gap is the lack of mention of any network or rate-limit behaviors, but this is minor for a read-only check.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides a full description for the single 'url' parameter ('The site to analyse, e.g. example.com or https://example.com'), giving 100% schema coverage. The tool description does not add additional parameter-level guidance beyond what the schema offers, so a baseline score of 3 is appropriate as per the calibration rules.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Checks') and resource ('whether AI crawlers are allowed to read a site'), enumerates the exact crawlers covered, and clearly differentiates itself from sibling tools like offpage_gaps, shopping, and entity_check by focusing on robots.txt-based AI visibility. This leaves no ambiguity about the tool's function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use it ('when someone asks why AI never mentions their site, or before any other AI-visibility work') and provides a contextual exclusion ('cannot detect blocking done at the CDN or firewall layer'). It does not name alternative tools, but the 'before any other AI-visibility work' phrasing implies this is a prerequisite rather than a competing option, which is sufficient guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
crawlbit_entity_checkBrand entity checkA
Checks whether a brand reads as a clear, consistent entity that AI can recognise: Organization schema, a consistent name, a description, a logo and linked profiles (sameAs). AI engines struggle to name a brand they cannot identify as a distinct thing. Use this when a brand is confused with another, or described vaguely, in AI answers. Measures the signals published on the site itself, not the brand's reputation across the web.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The site to analyse, e.g. example.com or https://example.com |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the tool measures on-site signals only and not web-wide reputation, which is a key behavioral trait. It could add more about return format or side effects, but for a read-only check this is sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, front-loaded with the purpose, and every sentence serves a purpose: purpose, context, usage, scope. No fluff or redundant repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity (one parameter, no output schema, no annotations), the description adequately covers what the tool does, when to use it, and its limitations. The only gap is not describing the output format, but that is not required for a straightforward check tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single url parameter, with a clear description in the schema. The tool description adds minimal semantics beyond implying the URL is the site to analyse. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool checks whether a brand reads as a clear, consistent AI-recognizable entity, listing specific signals (Organization schema, name, description, logo, sameAs). This is a specific verb+resource that clearly distinguishes it from sibling tools like offpage gaps or shopping.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit usage guidance is given: 'Use this when a brand is confused with another, or described vaguely, in AI answers.' It also excludes external reputation ('not the brand's reputation across the web'), clarifying its scope versus other possible tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
crawlbit_offpage_gapsOff-page gap finderA
Finds where a brand is missing OUTSIDE its own website: review platforms, directories, communities and 'best of' roundups relevant to its category. Off-page mentions on third-party pages are the strongest known predictor of AI citations, because engines weight what others say about a brand far above what the brand says about itself. Use this when a site is technically perfect yet still never cited. Returns the gaps and priorities; it does not publish anything or contact anyone.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The site to analyse, e.g. example.com or https://example.com |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly states that the tool is non-destructive ('does not publish anything or contact anyone') and describes what it returns ('gaps and priorities'). This adds useful context beyond the schema, though it could mention whether any crawling or network requests are made.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four sentences long, each earning its place: main action, rationale, when-to-use, and behavioral boundary. It is front-loaded with the core verb and resource, with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-parameter tool with no output schema, the description covers all essential decision-making information: what it finds, why it matters, when to use it, and what it returns. The high-level return phrase 'gaps and priorities' is adequate for an agent to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%: the single required `url` parameter already has a clear description and example. The tool description adds no additional parameter-level semantics, but the parameter is self-explanatory, so the schema alone is sufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with the specific verb 'Finds' and names concrete resource types (review platforms, directories, communities, 'best of' roundups), making it unmistakably distinct from sibling tools like shopping or crawler watch. It also clarifies the scope as 'OUTSIDE its own website,' further sharpening the purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit trigger: 'Use this when a site is technically perfect yet still never cited.' It also provides a when-not by stating that the tool 'does not publish anything or contact anyone,' so the agent knows it is only for analysis, not action.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
crawlbit_shoppingAI shopping readinessA
Checks whether AI shopping agents such as ChatGPT Shopping can find, understand and recommend a store's products: product schema completeness, price and availability signals, and the attributes those agents read. Use this for e-commerce sites, especially Shopify, when products never surface in AI recommendations. Inspects published structured data, so it cannot see a merchant feed submitted privately to a platform.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The site to analyse, e.g. example.com or https://example.com |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of disclosure. It clearly states the tool inspects published structured data and cannot see private feeds, which is a key limitation. It also names what it evaluates (schema completeness, price, availability, attributes), giving a solid behavioral overview.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the first states purpose, the second gives usage context, and the third describes a limitation. The description is front-loaded and concise without unnecessary fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description explains what the tool checks and its limitation, which sets appropriate expectations. It doesn't describe return format, but the scope of analysis is clear. For a diagnostic tool with one parameter and clear use cases, this is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter url is fully described in the schema with examples (example.com or https://example.com), achieving 100% schema coverage. The description adds no extra parameter-level detail beyond what the schema provides, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks whether AI shopping agents can find, understand, and recommend products, focusing on schema completeness, price/availability signals, and attributes. This specific verb+resource scope distinguishes it from sibling tools like offpage_gaps or crawler_watch.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to use for e-commerce sites, especially Shopify, when products never surface in AI recommendations. Also provides exclusion: it cannot see privately submitted merchant feeds, which helps avoid misuse. No explicit alternatives are named, but the when-to-use guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.
4 tool updates
v1.0.1- First observed
crawlbit_crawler_watch - First observed
crawlbit_entity_check - First observed
crawlbit_offpage_gaps - First observed
crawlbit_shopping
TDQS
Each tool targets a distinct aspect of AI visibility: off-page mentions, product schema for shopping agents, crawler access, and entity recognition. There is no overlap, so an agent can unambiguously select the right tool for a given diagnostic need.
All tools share the 'crawlbit_' prefix followed by a descriptive snake_case suffix (offpage_gaps, shopping, crawler_watch, entity_check). The naming convention is consistent and readable, with only a minor variation in compound-word structure that does not cause confusion.
Four tools is an ideal size for a focused diagnostic server. Each tool addresses a separate facet of AI visibility, keeping the scope tight without unnecessary bloat or overwhelming the agent.
The tool set covers the major pillars of AI visibility: being crawlable (crawler_watch), being recognizable (entity_check), having external validation (offpage_gaps), and being suitable for AI shopping (shopping). This is a complete diagnostic suite for the domain, with no obvious missing operations.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
Scan any website's AI readiness: AI search visibility and AI agent usability. Free, no auth.
Free AI-visibility and competitive Exposure Audit for any domain. No account, no API key.
Free SEO, GEO, and AEO audits: analyze any page or domain, AI-crawler access, agent readiness.
- VibeSEOOAuthdev.vibeseo
SEO research, audits, backlinks, GSC, and content workflow tools for AI agents.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEvaluates any website's AI visibility with 15 checks across crawlability, structure, content, and connectivity, and provides actionable fixes.10MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to perform instant SEO audits, check robots.txt, sitemaps, and AI crawler access for any URL without API keys.MIT
- AlicenseAqualityAmaintenanceProvides AI-visibility scoring and site auditing capabilities for websites, enabling agents to check how sites appear in AI engines like ChatGPT and Perplexity, run full SEO/security audits, and monitor changes over time.15405MIT
- AlicenseNot gradedqualityBmaintenanceEnables auditing AI search visibility: checks site readiness for AI crawlers and measures whether ChatGPT, Gemini, and Perplexity recommend your site, including verbatim answers and citation gap analysis.1374AGPL 3.0
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/amati032-dev/crawlbit-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server