ZapFetch MCP Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ZapFetch MCP ServerScrape the top post from news.ycombinator.com"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@zapfetchdev/mcp-server
MCP (Model Context Protocol) server for ZapFetch — APAC-native web scraping API for AI agents.
Use ZapFetch directly from Claude Desktop, Cursor, Windsurf, and any other MCP-compatible client.
Tools
zapfetch_scrape— scrape a single URLzapfetch_search— web search with optional content extractionzapfetch_crawl— crawl a website (async, returns job_id)zapfetch_crawl_status— poll crawl job progresszapfetch_map— discover URLs on a site (fast, no content)zapfetch_extract— extract structured data with a prompt + schemazapfetch_extract_status— poll extract job progress
Docs: https://docs.zapfetch.com
For docs lookups, Claude/Cursor/Windsurf can also use the auto-generated Mintlify MCP at https://docs.zapfetch.com/mcp.
Related MCP server: Kryfto
Prerequisites
Node.js 20+
A ZapFetch API key (get one here)
Install
npm install -g @zapfetchdev/mcp-serverConfigure
Claude Desktop
Edit ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):
{
"mcpServers": {
"zapfetch": {
"command": "npx",
"args": ["-y", "@zapfetchdev/mcp-server"],
"env": {
"ZAPFETCH_API_KEY": "zf-your-api-key"
}
}
}
}Cursor
Edit ~/.cursor/mcp.json:
{
"mcpServers": {
"zapfetch": {
"command": "npx",
"args": ["-y", "@zapfetchdev/mcp-server"],
"env": { "ZAPFETCH_API_KEY": "zf-your-api-key" }
}
}
}Windsurf
Edit ~/.codeium/windsurf/mcp_config.json:
{
"mcpServers": {
"zapfetch": {
"command": "npx",
"args": ["-y", "@zapfetchdev/mcp-server"],
"env": { "ZAPFETCH_API_KEY": "zf-your-api-key" }
}
}
}Install via Smithery
For a one-command install that auto-writes the config for your MCP client:
npx -y @smithery/cli install @zapfetchdev/mcp-server --client claude
# or --client cursor, --client windsurfSmithery will prompt for your ZapFetch API key once, then register the server in the right config file for you. This invokes the package's STDIO entry (bin.zapfetch-mcp → dist/index.js) — identical to the manual configs above.
Environment Variables
Variable | Required | Default | Description |
| STDIO | — | Your ZapFetch API key (STDIO mode only — HTTP mode takes the key per-request via |
| no |
| Override for self-host / dev |
| HTTP |
| HTTP server port (HTTP mode only) |
| Docker |
| Inside the Docker image only, switches between |
Usage Examples
After configuration, ask your AI assistant naturally — it will pick the right tool automatically.
Scrape a single page
"Scrape https://rakuten.co.jp and give me the main content as markdown."
Uses zapfetch_scrape. Best for a known URL where you want raw page content quickly. If the page is geo-blocked or returns sparse content, follow up with a search instead.
Search the web
"Find the top 5 recent blog posts about TypeScript 5.7 and summarize each one."
Uses zapfetch_search. Returns ranked results with optional content extraction. Useful when you don't have a specific URL yet, or as a fallback when a direct scrape comes up empty.
Crawl a site (multi-page)
"Crawl https://docs.example.com starting from the root, up to 50 pages, and summarize the authentication section."
Uses zapfetch_crawl to kick off an async job (returns a job_id), then zapfetch_crawl_status to poll until complete. The assistant handles the polling loop — you just wait for the result.
Tip: For large sites, map first (see below) to identify which URLs are worth crawling before committing.
Map a site (URL discovery)
"List all URLs under https://docs.example.com/api so I can decide which pages to scrape."
Uses zapfetch_map. Returns URLs only — no content fetched — so it's fast even on large sites. Pair with zapfetch_scrape to cherry-pick the pages you actually need:
"Map https://stripe.com/docs, then scrape the 3 pages most relevant to webhook setup."
Extract structured data
"Extract product name, price, currency, and stock status from these 5 rakuten.co.jp product URLs. Return as a JSON array."
Uses zapfetch_extract with a prompt and optional JSON schema. The job is async — zapfetch_extract_status polls it to completion. Good for turning arbitrary product pages, job listings, or articles into structured records at scale.
Poll extract job status
"Check whether the extract job job_abc123 is done."
Uses zapfetch_extract_status directly. You rarely need to ask for this by name — the assistant calls it automatically after zapfetch_extract — but it's useful if you started a job in a previous session and want to retrieve results later.
Combining tools
Tools compose naturally. A few common patterns:
Survey then scrape: map a large site to get all URLs, filter to the relevant ones, scrape each.
Search then scrape: search to find the canonical source for a topic, then scrape that page for full content.
Scrape with fallback: if
zapfetch_scrapereturns thin content (e.g. JS-heavy page), the assistant can fall back tozapfetch_searchto find a cached or mirror version.
Migrating from Firecrawl MCP
ZapFetch is Firecrawl-compatible at the API level, but this MCP uses zapfetch_* tool names (not firecrawl_*) to avoid conflicts if you run both. Capabilities are 1:1 — just update prompts referring to tool names.
HTTP transport (self-hosted)
For hosted / multi-tenant deployments, run the HTTP server instead of STDIO. Each request carries its own API key via Authorization: Bearer, so a single deployment can serve many users without sharing credentials.
Docker
docker run -d \
-p 3000:3000 \
-e ZAPFETCH_TRANSPORT=http \
docker.io/zapfetchdev/mcp:latestFrom npm
npm install -g @zapfetchdev/mcp-server
# then run the HTTP entry (never the STDIO one in this mode)
zapfetch-mcp-httpEndpoints
Path | Auth | Behavior |
| none | Returns |
|
| MCP JSON-RPC endpoint. Requests MUST include |
Example call
curl -X POST http://localhost:3000/mcp \
-H "Authorization: Bearer fc-YOUR-KEY" \
-H "Content-Type: application/json" \
-H "Accept: application/json, text/event-stream" \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'Security notes
Do not set
ZAPFETCH_API_KEYin HTTP mode. The server refuses to start (exit code 2) if the env var is present — this prevents a misconfigured container from silently serving every request from the operator's key.The HTTP server is stateless: each request builds its own transport + client, with the Bearer token flowing to the MCP tool handler via the SDK's native
extra.authInfo.tokenchannel. No cross-request state, no session leaks.Upstream ZapFetch API error strings are sanitized before transiting to HTTP clients — only a small allowlist of error codes (rate_limit, invalid_key, quota_exceeded, upstream_unavailable) passes through.
The stderr access log is strict-allowlist: only
ts / method / path / status / ms / origin_ip. Bearer tokens, bodies, and headers are never logged.
Development
pnpm install # or npm install
npm run typecheck
npm run build # -> dist/Local test with Claude Desktop pointed at your build:
{
"mcpServers": {
"zapfetch-dev": {
"command": "node",
"args": ["/absolute/path/to/zapfetch-mcp/dist/index.js"],
"env": { "ZAPFETCH_API_KEY": "zf-..." }
}
}
}License
MIT
Available Tools
7 toolszapfetch_crawlA
Crawl a website and extract content from multiple pages. Use this when the user wants to gather content from an entire site or section. Returns a job_id for async polling. Long-running — consider asking the user before invoking if the site is large. For a single page, use zapfetch_scrape. For URL discovery only (no content), use zapfetch_map.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The root URL to start crawling from | |
| limit | No | Maximum pages to crawl (default 50) | |
| maxDiscoveryDepth | No | Maximum link depth from root URL | |
| includePaths | No | Only crawl URLs matching these path patterns (regex) | |
| excludePaths | No | Skip URLs matching these path patterns (regex) | |
| crawlEntireDomain | No | Crawl entire domain, not just subpath of root URL |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, the description clearly states it is long-running and returns a job_id for async polling. It also advises considering user consent for large sites, which covers important behavioral aspects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Concise, with key information front-loaded. Every sentence adds value: defines the tool, specifies async nature, warns about runtime, and gives alternatives. No fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (crawling, async, multiple parameters), the description covers purpose, usage boundaries, behavioral notes, and alternatives. No output schema needed since it returns a job_id; the polling tool likely handles that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description does not repeat param details but adds context by explaining the overall behavior (async) and when to use. Since all parameters are well-documented in schema, the description adds value indirectly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it crawls a website and extracts content from multiple pages, using specific verbs ('crawl', 'extract'). It distinguishes from siblings: 'For a single page, use zapfetch_scrape. For URL discovery only (no content), use zapfetch_map.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells when to use (multi-page content gathering) and when not to (single page -> zapfetch_scrape, URL discovery -> zapfetch_map). Also warns about long-running nature and suggests asking user before invoking on large sites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zapfetch_crawl_statusA
Check the status of a running crawl job. Returns 'scraping' (in progress) or 'completed'/'failed'/'cancelled' (done). When status=completed, returns all crawled page content. Poll every 2-5 seconds until done.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | The crawl job ID returned from zapfetch_crawl |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the return states and behavior on completion, which adds value beyond schema. With no annotations provided, the description covers the key behavioral aspects well, though does not mention rate limits or side effects (likely none).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences that front-load the purpose, then detail outcomes, and end with actionable polling guidance. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description partially compensates by listing possible return states and content on completion. However, it does not specify the structure of the returned page content, leaving a minor gap. Overall adequate for a simple status poller.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single parameter job_id; the schema already describes it as the crawl job ID. The description adds no further semantic detail beyond referencing the source tool zapfetch_crawl, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks the status of a running crawl job, lists the possible statuses, and explains what happens when completed. It distinguishes itself from sibling tools like zapfetch_crawl by being the status-check counterpart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly advises polling every 2-5 seconds until done, which is a direct usage guideline. Also subtly indicates when to use this tool (after starting a crawl with zapfetch_crawl) by referencing the job_id from that tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zapfetch_extractA
Extract structured data from one or more URLs using natural language prompt + optional JSON schema. Use this when the user wants JSON-shaped data (not markdown), especially across multiple pages with the same structure (e.g. product listings, job postings, article metadata). Returns a job — poll with zapfetch_extract_status.
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | URLs to extract structured data from | |
| prompt | Yes | What to extract. Be specific — e.g. 'product name, price, and availability from each page' | |
| schema | No | Optional JSON schema describing the output shape. Improves consistency when extracting across many URLs. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It discloses that the tool returns a job which must be polled with a sibling tool. However, it does not mention any side effects, rate limits, or auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Concise two-sentence description plus additional parameter details in adjacent fields. Information is front-loaded (purpose, usage, behavior). Could be slightly more concise by merging the return value note into the first sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, description covers purpose, usage, return behavior, and parameter semantics adequately. The polling step is mentioned, which is critical. However, missing details on error handling or output format specifics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. Description adds value by explaining the optional schema field's purpose ('improves consistency when extracting across many URLs') but doesn't provide syntax examples.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses specific verb+resource ('Extract structured data from URLs') and clarifies the output format (JSON, not markdown). It explicitly distinguishes from sibling tools by mentioning the use case for structured data extraction from multiple pages.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly states when to use this tool ('when the user wants JSON-shaped data, especially across multiple pages with the same structure') and provides specific examples (product listings, job postings). Implicitly tells when not to use it (not for markdown).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zapfetch_extract_statusA
Check the status of a running extract job. Returns 'processing' (in progress) or 'completed'/'failed'/'cancelled' (done). When status=completed, returns the extracted structured data for all input URLs. Poll every 2-5 seconds until done.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | The extract job ID returned from zapfetch_extract |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden of behavioral disclosure. It correctly describes polling behavior and status outcomes, but does not mention rate limits, idempotency, or whether the job_id remains valid after completion. A 3 is appropriate as it covers key traits but lacks depth.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is highly concise: three sentences, no fluff. Every sentence adds value: first defines purpose, second details status values and data return, third gives polling guidance. Information is front-loaded with the most critical fact (status checking) first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has one parameter, no output schema, and no annotations, the description is nearly complete. It explains status values, data return on completion, and polling interval. Minor omission: does not clarify if the job_id can be reused or if status is final, but these are inferred from typical job models. A 4 is reasonable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% coverage with a description for 'job_id' that references the origin tool 'zapfetch_extract'. The description reinforces that job_id comes from the extract tool, adding context beyond the schema. No additional parameter information is needed since there's only one required parameter with a clear schema definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool checks the status of a running extract job, specifying the exact status values ('processing', 'completed', 'failed', 'cancelled') and what happens on completion. It uses specific verbs ('check', 'returns') and identifies the resource ('extract job status'), distinguishing it from sibling tools like 'zapfetch_extract' which presumably initiates the job.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to 'poll every 2-5 seconds until done', providing clear guidance on when and how to use this tool repeatedly. It implies usage after initiating an extract job with 'zapfetch_extract', and the polling interval is specified, helping the agent decide usage frequency without overloading the API.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zapfetch_mapA
Discover all URLs on a website without crawling content. Fast — use this to survey a site before deciding what to crawl. Returns list of URLs, optionally with titles/descriptions. Much cheaper than zapfetch_crawl (no content extraction). Use zapfetch_scrape or zapfetch_crawl on specific URLs from the result.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to map — discover all URLs on this site | |
| search | No | Filter discovered URLs by this search term (matches URL, title, or description) | |
| limit | No | Max URLs to return (default 100) | |
| sitemap | No | How to use the site's sitemap.xml: include (default), skip, or only |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. Discloses key behaviors: fast, no content extraction, returns URLs with optional metadata. Could add more detail on what 'map' means (e.g., follows links?), but covers the main behavioral traits adequately. No annotation contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: first explains what it does, second gives use case, third guides to alternatives. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, description doesn't detail return format, but states it returns list of URLs optionally with titles/descriptions. With 4 parameters and good schema coverage, the description is sufficiently complete for an agent to use correctly. A slight improvement would be mentioning that it follows links from the given URL to discover others.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and all parameters have descriptions. Description adds value by explaining the context of parameters: search filter matches URL, title, or description. The limit and sitemap are well-described in schema. Description reinforces the core purpose without adding redundant information, but overall parameter understanding is excellent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool discovers URLs without content, distinguishes from zapfetch_crawl (no content extraction). Uses specific verb 'Discover' and resource 'URLs on a website'. Siblings include zapfetch_crawl and zapfetch_scrape, and the description explicitly contrasts with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use: 'survey a site before deciding what to crawl'. Also tells when not to: 'use zapfetch_scrape or zapfetch_crawl on specific URLs from the result'. Names alternative tools and the decision flow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zapfetch_scrapeA
Scrape a single web page. Use this when the user wants to extract the content of ONE specific URL. Returns clean markdown (and other formats) of the page content. For crawling multiple URLs on a site, use zapfetch_crawl instead.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The URL to scrape | |
| formats | No | Output formats to return. Default: ['markdown'] | |
| onlyMainContent | No | Strip nav/footer/sidebar to keep main article only | |
| waitFor | No | Milliseconds to wait for JS rendering before extracting | |
| mobile | No | Emulate mobile user agent | |
| location | No | Geographic location — use for APAC sites needing local IPs |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries the burden. It states the tool returns 'clean markdown (and other formats)' but does not disclose potential limitations like JS rendering timeouts, error behavior, or rate limits. Adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no wasted words. Front-loaded with action and resource, then sibling differentiation. Perfectly concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-page scraping tool with a rich schema (6 parameters, all documented), the description is complete enough. The only gap is potential behavioral details like JavaScript rendering handling, but the schema's 'waitFor' parameter covers that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds no extra meaning beyond the schema's parameter descriptions, but the schema already has good descriptions for each parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states it scrapes a single web page and returns clean markdown (and other formats). It clearly distinguishes from sibling zapfetch_crawl by explicitly saying 'for crawling multiple URLs on a site, use zapfetch_crawl instead'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use this when the user wants to extract the content of ONE specific URL' and names the alternative for crawling. Provides clear when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zapfetch_searchA
Search the web and optionally scrape results. Use this when the user wants to find information across multiple sites — like a programmable Google. Returns structured list of result URLs with titles and snippets. Use zapfetch_scrape to fetch full content of a specific result.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | The search query | |
| limit | No | Max results to return (default 10) | |
| tbs | No | Time-based search filter, e.g. 'qdr:d' for past day, 'qdr:w' for past week | |
| location | No | Search location hint, e.g. 'Tokyo, Japan' for APAC-focused results |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the full burden. The description states that the tool 'optionally scrapes results' and returns a 'structured list of result URLs with titles and snippets.' This provides some behavioral context, but does not disclose details like whether the operation is read-only, if there are any rate limits or delays, or what happens if the query returns no results. With zero annotations, more behavioral context would be beneficial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded: the first sentence states the core purpose, and subsequent sentences add context and sibling differentiation. Every sentence adds value, and the entire description fits in two sentences with a third for the alternative tool. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there is no output schema, the description covers the return format (structured list of result URLs with titles and snippets). It clearly explains what the tool does and how it relates to its sibling. It could potentially mention that this tool is designed for multi-site search and that scraping is optional, but it already implies that. For a search tool with clear sibling comparisons, it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage, meaning each parameter already has a description. The tool description does not add any additional meaning beyond what the schema provides. According to the rubric, when schema_description_coverage is high (>80%), the baseline is 3 even with no param info in the description. The description does not explain how parameters interact or provide examples, but the schema already covers their basic meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool searches the web and optionally scrapes results. It uses specific verbs ('Search' and 'scrape') and identifies the resource ('the web'). It differentiates from sibling zapfetch_scrape by noting that this tool returns result URLs with titles and snippets, while zapfetch_scrape fetches full content of a specific result.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool ('when the user wants to find information across multiple sites') and names an alternative tool to use when you need full content of a specific result ('use zapfetch_scrape'). This provides clear differentiation from its primary sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.0- First observed
zapfetch_crawl - First observed
zapfetch_crawl_status - First observed
zapfetch_extract - First observed
zapfetch_extract_status - First observed
zapfetch_map - First observed
zapfetch_scrape - First observed
zapfetch_search
TDQS
Scored across 7 tools
Each tool has a clearly distinct purpose: crawling multiple pages (zapfetch_crawl), checking crawl status (zapfetch_crawl_status), structured extraction (zapfetch_extract, zapfetch_extract_status), URL discovery (zapfetch_map), single-page scraping (zapfetch_scrape), and web search (zapfetch_search). Descriptions explicitly contrast them, eliminating ambiguity.
Tool names follow a consistent prefix 'zapfetch_' with verb_noun pattern for actions (crawl, crawl_status, extract, extract_status, map, scrape, search). Minor inconsistency: 'crawl' vs 'crawl_status' and 'extract' vs 'extract_status' use a suffix, while others are single verbs. Still highly predictable.
Seven tools cover the core web scraping and search functionality without redundancy. Each tool serves a unique purpose, and the count is appropriate for the domain. Not too few nor too many.
The tool set covers a complete workflow: discover URLs (zapfetch_map), scrape single pages (zapfetch_scrape), crawl multiple pages (zapfetch_crawl), extract structured data (zapfetch_extract), search the web (zapfetch_search), and poll async jobs (zapfetch_crawl_status, zapfetch_extract_status). No obvious gaps for common web scraping tasks.
Maintenance
Related MCP Connectors
Scrape, crawl and search the web for AI agents via MCP.
Cloud scraping & crawling API for AI agents. Turn any URL into clean, LLM-ready markdown.
Web scraping for agents. Point it at a URL and it returns the page as clean markdown, JavaScript-rendered pages included. Point it at a site and it maps the URLs or crawls the section you need in the background, a few pages at a time so results fit in the conversation. Search the web and read full pages, extract fields with a JSON schema you define (validated, never invented), read a store's catalogue or a blog's posts from the platform's own feed, and check whether a page has changed. Failed requests cost nothing. The free plan includes 1,500 credits a month.
Direct access to 60+ scraping and search tools. Extract structured data from Google (Search, Maps, Trends), Amazon, Airbnb, Social Media, and any web page directly into your AI agent.
Related MCP Servers
- AlicenseAqualityDmaintenanceMCP-native web scraping and search API for AI agents. Converts any URL to clean Markdown with 90% success rate, including Cloudflare-protected sites and JS SPAs. Real-time web search via Brave Search API. CAPTCHA solving built-in. 10 free scrapes/day.55 npm5MIT
- FlicenseNot gradedqualityCmaintenanceProvides 42+ MCP tools for browser automation, web scraping, and search, enabling AI agents like Claude and Cursor to browse, extract data, and run research agents on the live web.9-
- -licenseNot gradedqualityCmaintenanceSelf-hosted MCP server that provides web scraping and crawling tools, integrating seamlessly with AI frameworks like OpenAI Agents SDK, Cursor, and Claude Code.4-
- AlicenseBqualityBmaintenanceMCP server exposing 35 web scraping and SERP tools for general scraping, Google services, e-commerce sites, social media, and search engines. Enables MCP clients like Claude to scrape web pages and search results via natural language.3538 npmMIT