"Real-time web scraping tools" matching MCP connectors:
Matching Connector Tools:
Deterministic AI agent microtools, no accounts/API keys. fetch_extract: 98% token cut. 38 tools.
The only News based AI MCP your agents will ever need — custom categories, global regions, and time-scoped results in one tool. We use multi-vector & sparse-hybrid search to search through thousands of articles across the world to find the exact news you're looking for.
AI-first file sharing and collaboration. 251 tools give agents a full workspace: file storage, branded shares, comments, workflows, and built-in RAG. 50GB free, no credit card.
**ColdState Knowledge Search MCP Server** https://github.com/daniel-coldstate/coldstate-mcp Semantic search over 64.6M knowledge entries — the structured alternative to web search APIs and web scraping for LLM agents. No crawling, no rate limits, sub-3s responses. Cloud-hosted at services.coldstate.ai
Collaborative, cache-first web search for agents — cited answers from a shared live-web pool.
Diffbot MCP — Knowledge Graph company enrichment + web content extraction (diffbot.com)
Multilingual YouTube → Knowledge Pack engine. Paste a video URL and get a structured pack — summary, key ideas, glossary, quiz, transcript with timestamps — in Spanish, Portuguese, German, or English. Anonymous endpoint plus OAuth-gated tools for library search, RAG Q&A on a single pack, and Anki export.
Semantic search over a 200-chunk ACLM lifestyle medicine knowledge base.
Brave Search MCP — independent web index (no Google/Bing dependency)
Direct access to 40+ scraping and search tools. Extract structured data from Google (Search, Maps, Trends), Amazon, Airbnb, Social Media, and any web page directly into your AI agent.
Shared knowledge base for AI agents. Semantic search across agents, no setup required — just a URL.
100+ MCP tools for AI agents: content metadata, trade intelligence, business-expertise analysis.
Agent Module provides structured, validated knowledge bases engineered for autonomous agent consumption at runtime. Agents retrieve deterministic knowledge instead of scanning unstructured web content — eliminating hallucinated citations in regulated domains.
Knowledge Base von designare.at – Michael Kanda, Web & KI aus Wien. Semantische Suche über RAG.
AI web extraction: send URLs + a JSON Schema, get clean structured data. Pay-per-use via x402.
VaultCrux Platform — 60 tools: retrieval, proof, intel, economy, watch, org
Real-time web search with answer-ready results for Claude, Cursor and any MCP client. A Tavily alternative: same speed, 20.2% fewer tokens, higher answer quality (60.7% of decided duels won) on a public benchmark. Hosted on mcp.serpdive.com or npx serpdive-mcp.
Unstructured document processing for LLM pipelines. Upload as PDF/DOCX/TXT any supported files, extract structured data (PII-redacted), build LLM-ready datasets, and search/export results — all via MCP tools (document.process, job.status, job.result, dataset.build, dataset.search, dataset.export).