webdatatools-rag-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| APIFY_TOKEN | Yes | Your Apify API token. This server uses your own Apify API token; every tool call runs a WebDataTools Actor under your Apify account and is billed to your Apify credit. Get a free token at https://console.apify.com/settings/integrations. |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| ai_web_searchA | AI Web Search runs a Google search, fetches the top organic results and returns clean Markdown per result — one call turns a question into LLM-ready context for agents, RAG and MCP. Billed to your own Apify account: ~$0.005 per result (Apify free-plan price, lower on paid plans). |
| website_to_markdownB | Content Crawler turns any site into clean Markdown per page for LLMs, RAG pipelines and vector DBs — no headless browser, $1 per 1,000 pages. Billed to your own Apify account: ~$0.001 per result (Apify free-plan price, lower on paid plans). |
| article_extractorA | Article & News Extractor returns clean article text, title, author(s), publish/modified date, tags and images from any news or blog URL — one row per URL, as Markdown, plain text or HTML. Billed to your own Apify account: ~$0.002 per result (Apify free-plan price, lower on paid plans). |
| structured_data_extractorA | Structured Data & JSON-LD Extractor reads every Schema.org JSON-LD block and Open Graph tag on a page and returns one clean row per URL — product, price, rating, article, job posting, event, recipe, FAQ and breadcrumb data, ready for RAG pipelines and SEO rich-result audits. Billed to your own Apify account: ~$0.002 per result (Apify free-plan price, lower on paid plans). |
| google_news_scraperA | Google News Scraper returns headlines from Google News RSS search and topic feeds for any keyword, site: or when: query, and resolves each article's real publisher URL — one row per article. Billed to your own Apify account: ~$0.0005 per article (Apify free-plan price, lower on paid plans). |
| press_release_monitorA | Press Release Monitor returns keyword-matched press releases from PR Newswire, Business Wire and GlobeNewswire RSS feeds — company, publish date, category and summary, one row per release. Billed to your own Apify account: ~$0.001 per result (Apify free-plan price, lower on paid plans). |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 6 tools
Most tools are clearly differentiated by source or output type: general web search, site-to-Markdown crawl, article extraction, structured data, Google News, and press releases. The main overlap is between website_to_markdown and article_extractor when processing a single news/blog article, but their broader purposes are still distinguishable.
All names use snake_case and are descriptive, which makes them readable and predictable. The patterns vary slightly (action-oriented names like ai_web_search vs. noun-extractor names like article_extractor), but the set remains consistent overall.
Six tools is a well-scoped size for a web-data RAG server. Each tool has a distinct source or extraction focus, so the surface does not feel bloated or thin.
The set covers core web-data acquisition workflows: search, crawling, article extraction, structured data, news, and press releases. Minor gaps such as PDF extraction, social media sources, or advanced crawl controls exist, but the core RAG pipeline needs are well represented.