webdatatools-rag-mcp
Provides AI web search capabilities using Google results, returning clean page text and Markdown for LLM context and RAG pipelines.
Provides Google News scraping and monitoring via RSS search by keyword, topic, or site, returning news articles for AI agents.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@webdatatools-rag-mcpcrawl example.com and convert it to Markdown for my RAG pipeline"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
WebDataTools Web content for AI & RAG MCP server
webdatatools-rag-mcp
An MCP server with 6 web content for ai & rag tools for AI agents — Claude Desktop, Cursor, Cline or any MCP client. Web search with page text, website-to-Markdown crawling, clean article extraction, JSON-LD/structured data, Google News and press-release monitoring — ready for LLM context and RAG pipelines.
This server uses your own Apify API token. Every tool call runs a WebDataTools Actor under your Apify account and is billed to your Apify credit — pay per result, the price is in each tool description. Your token is only sent to Apify's API.
Quick start
Requires Node.js 18+.
APIFY_TOKEN=apify_api_... npx -y github:paulet4a-commits/webdatatools-rag-mcpGet a free token (the free plan includes monthly credit): https://console.apify.com/settings/integrations
Related MCP server: websearch-skill
Claude Desktop / Cursor
Add this to claude_desktop_config.json (Claude Desktop) or .cursor/mcp.json (Cursor):
{
"mcpServers": {
"webdatatools-rag": {
"command": "npx",
"args": [
"-y",
"github:paulet4a-commits/webdatatools-rag-mcp"
],
"env": {
"APIFY_TOKEN": "apify_api_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"
}
}
}
}Tools (6)
Tool | What it does | Price (free plan) | Backing Actor |
| AI Web Search & Read: Google results as clean Markdown | $0.005 / result | |
| Website to Markdown — Content Crawler for LLM & RAG | $0.001 / result | |
| Article & News Extractor (clean text, author, date, markdown) | $0.002 / result | |
| Structured Data & JSON-LD Extractor (Schema.org, Open Graph) | $0.002 / result | |
| Google News Scraper (RSS search by keyword, topic, site) | $0.0005 / article | |
| Press Release Monitor: PR Newswire, BusinessWire, GlobeNewswire | $0.001 / result |
More WebDataTools MCP servers
webdatatools-mcp-server — the 10 most popular tools in one server
webdatatools-domain-mcp — Domain & website intelligence
webdatatools-social-mcp — Search, video & social data
webdatatools-leads-mcp — Leads, jobs & company data
webdatatools-dev-mcp — Developer, app & research data
License
MIT
Available Tools
6 toolsai_web_searchA
AI Web Search runs a Google search, fetches the top organic results and returns clean Markdown per result — one call turns a question into LLM-ready context for agents, RAG and MCP. Billed to your own Apify account: ~$0.005 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| urls | No | URLs to read directly (skip search) — Enter specific page URLs to fetch and convert to Markdown, e.g. https://docs.apify.com/platform. When set, no Google search is performed at all — this becomes a pure URL-to-Markdown reader. | |
| query | No | Search query — Enter a single question or search phrase, e.g. what is web scraping. Ignored if "queries" or "urls" is also set. One SERP is fetched and the top organic results are read and returned as Markdown. Example: "what is web scraping". | |
| queries | No | Search queries (batch) — Enter multiple search queries to run in one call, e.g. best crm for startups. Overrides "query" when non-empty. Leave empty to use "query" instead. Example: ["best crm for startups"]. | |
| maxResults | No | Max results per query — Enter how many organic results to read per query, e.g. 3. Each result is one billed dataset row, so 3 results costs about 3x the per-result price. | |
| countryCode | No | Country code (gl) — Enter the 2-letter country code Google should localise results for, e.g. us, gb, de. | us |
| languageCode | No | Language code (hl) — Enter the 2-letter interface language code, e.g. en, es, fr. | en |
| outputFormat | No | Output format — Choose markdown (clean Markdown, best for LLMs), text (plain text) or both. Options: markdown = Markdown; text = Plain text; both = Both. | markdown |
| maxCharsPerResult | No | Max characters per result — Enter the maximum characters to keep per result page, e.g. 8000. Long pages are truncated (truncated:true) to keep the response inside your LLM's context window. | |
| includeSnippetOnly | No | Snippet-only (cheap mode, no page fetch) — Turn this on to return only the SERP title/url/snippet for each result without fetching and reading the page — much faster and works even without page-fetch access, but markdown/text come back null. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does useful work: it discloses that billing hits the caller's own Apify account at ~$0.005 per result, and that output is Markdown per result. It omits auth/prerequisite details, rate limits, and failure modes, so it is strong but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The functional core (search → fetch → Markdown) is front-loaded, and the pricing detail follows as supporting context. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a nine-parameter tool with no annotations and no output schema, the description conveys the return shape (clean Markdown per result) and the cost model, which are the key things an agent needs. It is nearly complete, missing only operational details like auth setup and rate-limit behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all nine parameters are already richly documented in the schema (including mode exclusivity and billing implications of maxResults). The description adds no parameter-level meaning beyond that, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb+resource chain: runs a Google search, fetches top organic results, returns clean Markdown per result. This clearly distinguishes it from siblings like website_to_markdown or article_extractor, which do not perform search. An agent can identify the tool's function without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description frames the use case (turning a question into LLM-ready context for agents, RAG, MCP) but never states when to prefer it over the overlapping siblings website_to_markdown or article_extractor, nor explicit exclusions. Usage mode selection (search vs direct URL read) is documented in the schema, not the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
article_extractorA
Article & News Extractor returns clean article text, title, author(s), publish/modified date, tags and images from any news or blog URL — one row per URL, as Markdown, plain text or HTML. Billed to your own Apify account: ~$0.002 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | Article URLs — Enter the article or blog post URLs to extract, one row is returned per URL, e.g. https://blog.apify.com/best-web-scraping-tools/. Works on news sites, blogs and any page that publishes a Schema.org Article/NewsArticle JSON-LD block, Open Graph tags, or a plain readable body. Example: ["https://blog.apify.com/best-web-scraping-tools/"]. | |
| outputFormat | No | Output format — Choose which body format(s) to include in each row. "Markdown" is the smallest and best for LLM/RAG ingestion; "all" returns markdown, text and html together for debugging or comparison. Options: markdown = Markdown; text = Plain text; html = HTML; all = All (markdown + text + html). | markdown |
| includeImages | No | Include images — Keep this on to return the mainImage and images fields and keep image references in the markdown/html body. Turn it off for a smaller, text-only dataset. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it discloses meaningful traits: one row per URL, selectable output formats, and the cost model ('~$0.002 per result', billed to the caller's own Apify account). It stops short of covering failure modes for unsupported pages, rate limits, or the auth/token requirement implied by 'your own Apify account'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Roughly two sentences: the capabilities/output list is front-loaded, followed by the pricing caveat. Dense but nearly waste-free; the dollar figure is arguably useful rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter extraction tool with no output schema and no annotations, the description supplies the essential missing context: what fields come back, one-row-per-URL granularity, format options, and cost. It omits error/unsupported-page behavior and auth prerequisites, which keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces that output is one row per URL and that Markdown/plain text/HTML are selectable, but adds no format syntax, enum nuance, or default behavior beyond what the schema already documents for url, outputFormat, and includeImages.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource ('Article & News Extractor returns clean article text, title, author(s), publish/modified date, tags and images from any news or blog URL') and enumerates the exact output fields, so the agent knows precisely what it produces. It does not, however, explicitly differentiate itself from siblings like website_to_markdown or structured_data_extractor, which it partly overlaps with, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the phrase 'from any news or blog URL' — there is no explicit statement of when to prefer this over website_to_markdown or structured_data_extractor, nor any when-not conditions. The billing note gives some practical context but is not routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
google_news_scraperA
Google News Scraper returns headlines from Google News RSS search and topic feeds for any keyword, site: or when: query, and resolves each article's real publisher URL — one row per article. Billed to your own Apify account: ~$0.0005 per article (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| topics | No | Section topics — Enter Google News section topics to fetch instead of (or in addition to) a search, e.g. TECHNOLOGY or BUSINESS. Accepted values: WORLD, NATION, BUSINESS, TECHNOLOGY, ENTERTAINMENT, SCIENCE, SPORTS, HEALTH. Unknown values are ignored. | |
| country | No | Country — Enter the 2-letter country/edition code, e.g. US for the United States or GB for the United Kingdom. Combined with Language to build the hl/gl/ceid feed parameters. | US |
| queries | Yes | Search queries — Enter Google News search terms, one row is returned per matching article. Supports Google's search operators, e.g. web scraping, "exact phrase", site:reuters.com AI, or -unwanted. Leave empty and use Topics below instead if you only want section feeds. Example: ["web scraping"]. | |
| language | No | Language — Enter the 2-letter interface language code Google News should use, e.g. en for English or fr for French. Combined with Country to build the hl/gl/ceid feed parameters. | en |
| sinceDays | No | Only articles from the last N days — Enter how many days back to search, e.g. 7 for the last week. Appends Google's when:Nd search operator to every query (topic feeds ignore this — they are always "latest"). Enter 0 to disable and return whatever Google's default ranking gives. | |
| maxItemsPerQuery | No | Max items per query/topic — Enter the maximum number of articles to keep per query or topic, e.g. 50. Google's own feed rarely returns more than ~100 items for a single request no matter how high this is set. | |
| resolvePublisherUrls | No | Resolve publisher URLs — Keep this on to follow each article's news.google.com redirect link and fetch the publisher's real URL (2 extra requests per article, using an undocumented Google endpoint — best-effort, see README). Turn it off for a much faster run that only returns the Google News link. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose real behavioral traits: billing to the user's own Apify account at ~$0.0005/article, that resolvePublisherUrls costs 2 extra requests per article via an undocumented best-effort endpoint, and that topic feeds ignore sinceDays. It does not cover rate limits or failure behavior beyond the 'best-effort' note, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences: purpose first, then cost and row granularity. No filler, and pricing is placed where a cost-sensitive agent will see it. Slightly overstuffed with parenthetical pricing detail, but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter scraper with no output schema and no annotations, the description covers what matters: row-per-article output shape, cost model, and the performance/accuracy tradeoff of URL resolution. It leaves some return-field detail unstated, but the essential selection and invocation context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 7 parameters in detail (operator support, country/language feed construction, bounds). The description adds no per-parameter syntax beyond reinforcing the query types, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('returns headlines from Google News RSS search and topic feeds') and adds a concrete scope detail ('one row per article'). It implicitly separates itself from generic siblings like ai_web_search by narrowing to Google News RSS, but it never names an alternative tool, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It indicates supported query forms (keyword, site:, when:) and notes that topic feeds are an alternative to searching, which implies usage. However, it never states when to pick this tool over ai_web_search, article_extractor, or press_release_monitor, and there are no explicit exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
press_release_monitorA
Press Release Monitor returns keyword-matched press releases from PR Newswire, Business Wire and GlobeNewswire RSS feeds — company, publish date, category and summary, one row per release. Billed to your own Apify account: ~$0.001 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| sources | No | Wire sources — Select which press-release wires to read, e.g. prnewswire. Leave empty to read all three. Each source is fetched as its own public RSS feed, so a wire that is briefly unavailable produces one error row instead of failing the whole run. Options: prnewswire = PR Newswire; businesswire = Business Wire; globenewswire = GlobeNewswire. | |
| keywords | No | Keywords — Enter keywords to match against each release's title and summary, e.g. funding, acquisition. Case-insensitive substring match; a release needs only one hit to pass. Leave empty to keep every release from the selected wires. | |
| sinceHours | No | Only releases from the last N hours — Enter how many hours back to keep releases, e.g. 24. Releases without a parseable publish date are dropped when this is set. Leave at 0 to disable the time filter and keep every release currently in the feed. | |
| includeBody | No | Fetch full release text — Turn this on to fetch each release's page and add up to 8,000 characters of plain-text body. Only works for PR Newswire and GlobeNewswire — Business Wire's own website blocks non-browser requests, so its body is always null. | |
| maxItemsPerSource | No | Max items per source — Enter the most releases to keep per wire after filtering, e.g. 100. Each wire's RSS feed only ever contains its latest 20-50 releases, so this mostly matters when you run on a tight schedule and want to cap cost. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers valuable operational context: billing to the caller's own Apify account, a per-result price (~$0.001, lower on paid plans), and one-row-per-release output semantics. It stops short of stating rate limits, pagination, or auth setup, but the cost/billing disclosure is genuinely additive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the core capability is front-loaded before the pricing note. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, zero-required-param scraper with no output schema, the description gives sufficient scope, returned fields, and cost expectations. Since no annotations exist, it would be slightly stronger with an explicit read-only/no-mutation statement, but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each of the five parameters is already fully documented in the schema (including the enum titles, error-row behavior, and the Business Wire body caveat). The description adds no parameter-level syntax or format detail, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (returns keyword-matched press releases), names the exact three source wires, and enumerates the returned fields (company, publish date, category, summary). An agent can distinguish this from google_news_scraper and the other siblings purely from the description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The text is purely descriptive of behavior and pricing; it never says when to choose this over google_news_scraper or the other sibling scrapers, nor any when-not condition. Usage is only faintly implied by the tool name, so guidance is effectively absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
structured_data_extractorA
Structured Data & JSON-LD Extractor reads every Schema.org JSON-LD block and Open Graph tag on a page and returns one clean row per URL — product, price, rating, article, job posting, event, recipe, FAQ and breadcrumb data, ready for RAG pipelines and SEO rich-result audits. Billed to your own Apify account: ~$0.002 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | URLs — Enter the page URLs to extract structured data from, one row is returned per URL, e.g. https://www.allbirds.com/products/mens-tree-dashers. Works on product, article, job, event, recipe and FAQ pages — anything that publishes Schema.org JSON-LD or Open Graph tags. Example: ["https://www.apify.com"]. | |
| followRedirects | No | Follow redirects — Keep this on to follow HTTP redirects to the final page, e.g. a shortened or tracking URL. Turn it off to get an error row instead when a URL 301/302s, useful for auditing which URLs redirect. | |
| includeOpenGraph | No | Include Open Graph / Twitter Card data — Keep this on to return the openGraph field (og:title, og:description, og:image, og:type, og:site_name, twitter:card and related tags). Most pages publish these even without JSON-LD. | |
| includeRawJsonLd | No | Include raw JSON-LD — Keep this on to also return the page's raw parsed JSON-LD documents in the jsonLd field (capped at 400 KB per row), useful when you need a schema type the mapped fields do not cover. Turn it off for a smaller, cheaper-to-store dataset. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and delivers non-obvious operational facts: billing is charged to the caller's own Apify account at roughly $0.002 per result. It omits run-time behavior such as auth setup, rate limits, and pagination, but the cost/billing disclosure is genuine value beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action before the cost note. The first sentence is dense but every clause lists concrete extracted types rather than filler; the pricing sentence is short and earns its place as actionable cost context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description must convey the return shape — and it does ('one clean row per URL' with the enumerated field categories). For a read-only scraping tool with fully documented parameters, this is nearly complete; only the absence of return-format details like the row field names holds it below a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (urls, followRedirects, includeOpenGraph, includeRawJsonLd) are already documented in detail. The description adds only the 'one row per URL' output-shape hint and no parameter-level syntax or format detail, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('reads every Schema.org JSON-LD block and Open Graph tag on a page') and enumerates the payload it returns (product, price, rating, article, job posting, event, recipe, FAQ, breadcrumb). An agent can distinguish it from siblings like article_extractor or website_to_markdown without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context is supplied — 'ready for RAG pipelines and SEO rich-result audits' tells the agent the intended scenarios. However, it never names an alternative tool or states a when-not-to-use condition (e.g., when to pick article_extractor instead), so it falls short of explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
website_to_markdownB
Content Crawler turns any site into clean Markdown per page for LLMs, RAG pipelines and vector DBs — no headless browser, $1 per 1,000 pages. Billed to your own Apify account: ~$0.001 per result (Apify free-plan price, lower on paid plans).
| Name | Required | Description | Default |
|---|---|---|---|
| maxDepth | No | Max link depth — Enter how many links deep to follow from a start URL, e.g. 3. Set 0 to crawl only the start URL(s). | |
| maxPages | No | Max pages — Enter the maximum number of pages to crawl and convert, e.g. 50. This is the billed unit ($1 per 1,000 pages) — the crawl stops exactly at this count. | |
| startUrls | Yes | Start URLs — Enter the page(s) to start crawling from, e.g. https://docs.apify.com/platform. The crawler follows links from here in breadth-first order. Example: [{"url":"https://docs.apify.com/platform"}]. | |
| useSitemap | No | Seed from sitemap.xml — Turn this on to also read /sitemap.xml on each start URL's domain and add its URLs (filtered by the settings above) to the crawl queue, up to Max pages. | |
| outputFormat | No | Output format — Choose what content to put in each row: Markdown only, plain text only, or both. Markdown preserves headings, lists, code blocks and tables. Options: markdown = Markdown only; text = Plain text only; both = Markdown and plain text. | markdown |
| respectRobots | No | Respect robots.txt — Turn this on to skip URLs disallowed by the site's robots.txt file (recommended and on by default). | |
| sameDomainOnly | No | Same domain only — Turn this on to only follow links on the same domain as the start URL (www. is treated as the same domain), and off to also follow links to other domains. | |
| removeSelectors | No | Extra CSS selectors to remove — Optional. Extra CSS selectors to strip before extracting content, e.g. .cookie-banner or #newsletter-signup, on top of the built-in nav/header/footer/aside removal. | |
| excludePathPatterns | No | Exclude URL patterns — Optional. Regular expressions tested against the full URL; a match is skipped, e.g. \.(png|jpe?g)$ to skip images. Defaults cover binary files and login/signup pages. | |
| includePathPrefixes | No | Include path prefixes — Optional. Only crawl URLs whose path starts with one of these prefixes, e.g. /docs. Leave empty to crawl every path allowed by the other settings. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses the cost model ('$1 per 1,000 pages', billed to the caller's own Apify account) and the no-headless-browser characteristic, but says nothing about credentials required, crawl duration, throttling, or what happens on failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Only two sentences, but both are promotional and lead with the brand name rather than the operation. The detailed pricing in sentence two duplicates the maxPages schema description, so part of the text does not earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter crawl tool with no annotations and no output schema, the description gives the gist of the result ('clean Markdown per page') and the cost, but omits auth/account requirements, per-result row shape, and run mechanics. It is adequate but leaves real gaps given the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all ten parameters thoroughly, including defaults, bounds and examples. The description adds no parameter-level meaning beyond the schema — its pricing line duplicates what maxPages already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete verb and resource ('turns any site into clean Markdown per page'), so the core operation is clear despite the marketing framing around 'Content Crawler'. It does not differentiate the tool from siblings like article_extractor or structured_data_extractor, which also extract content from pages.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied via the audience callout ('for LLMs, RAG pipelines and vector DBs'), which tells the agent the broad context but not when to prefer this over article_extractor or structured_data_extractor. No explicit when-not conditions or alternative routing are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
ai_web_search - First observed
article_extractor - First observed
google_news_scraper - First observed
press_release_monitor - First observed
structured_data_extractor - First observed
website_to_markdown
TDQS
Scored across 6 tools
Most tools are clearly differentiated by source or output type: general web search, site-to-Markdown crawl, article extraction, structured data, Google News, and press releases. The main overlap is between website_to_markdown and article_extractor when processing a single news/blog article, but their broader purposes are still distinguishable.
All names use snake_case and are descriptive, which makes them readable and predictable. The patterns vary slightly (action-oriented names like ai_web_search vs. noun-extractor names like article_extractor), but the set remains consistent overall.
Six tools is a well-scoped size for a web-data RAG server. Each tool has a distinct source or extraction focus, so the surface does not feel bloated or thin.
The set covers core web-data acquisition workflows: search, crawling, article extraction, structured data, news, and press releases. Minor gaps such as PDF extraction, social media sources, or advanced crawl controls exist, but the core RAG pipeline needs are well represented.
Maintenance
Related MCP Connectors
Scrape, crawl and search the web for AI agents via MCP.
Web MCP: scrape/crawl sites, web search, brand assets, app stores, YouTube, Reddit, Hacker News.
Live AI-native web search with citations. One tool for every MCP client. Flat per-request pricing.
Fetch pages as markdown, search web and news, extract structured data. For AI agents.
Related MCP Servers
- AlicenseAqualityDmaintenanceMCP-native web scraping and search API for AI agents. Converts any URL to clean Markdown with 90% success rate, including Cloudflare-protected sites and JS SPAs. Real-time web search via Brave Search API. CAPTCHA solving built-in. 10 free scrapes/day.55 npm5MIT
- AlicenseAqualityBmaintenanceEnables AI agents to perform multi-engine web search, fetch web pages, and extract clean Markdown content via MCP, with no API keys required.340 PyPI8MIT
- AlicenseAqualityCmaintenanceGives MCP-capable agents live web access: search the web, scrape pages into Markdown (including JavaScript-heavy and bot-protected sites), and extract named fields as JSON, with job polling, token-aware content offloading, and built-in research guidance. Ships as a self-hostable stdio or HTTP service with spend caps and per-request key support.7MIT
- AlicenseAqualityBmaintenanceThis MCP server equips AI agents with ten web-data tools for web search, page reading, site crawling, contact and tech-stack detection, company profiling, e-mail/DNS security checks, e-mail validation, package health, and raw Google search. Each call runs Actors on the user's own Apify account, returning trimmed, readable dataset rows.10MIT