"Information about extraction techniques or processes" matching MCP connectors:
Matching Connector Tools:
For queries a model can't confidently place: resolves or declines. Built to decline, not guess.
Extract PDFs to Markdown, RAG chunks and cited tables; publish tracked Doc Links with read stats.
Parse PDF/Word/PPT/HTML to Markdown; tables as JSON, image extraction, RAG chunking, page ranges.
Diffbot MCP — Knowledge Graph company enrichment + web content extraction (diffbot.com)
Multilingual YouTube → Knowledge Pack engine. Paste a video URL and get a structured pack — summary, key ideas, glossary, quiz, transcript with timestamps — in Spanish, Portuguese, German, or English. Anonymous endpoint plus OAuth-gated tools for library search, RAG Q&A on a single pack, and Anki export.
Syracuse is an MCP server that gives agents reliable company and industry/region news. Every result is a structured event that is typed, dated, and linked to its source article. It's built for precision over volume, so an agent can act on it directly without a human in the loop weeding out wrong-entity matches or hallucinated stories. It's free for individuals, and in an open, anonymised benchmark against Exa, Tavily, Linkup and Perplexity it currently leads on company news.
AI web extraction: send URLs + a JSON Schema, get clean structured data. Pay-per-use via x402.
Connect your AI to any database — PostgreSQL, MySQL, or SQL Server — in seconds.
Real-time web search with answer-ready results for Claude, Cursor and any MCP client. A Tavily alternative: same speed, 20.2% fewer tokens, higher answer quality (60.7% of decided duels won) on a public benchmark. Hosted on mcp.serpdive.com or npx serpdive-mcp.
Persistent memory for AI assistants. Save once; recall from Claude, ChatGPT, or any MCP client.
LLMtoMD is the memory layer for AI coding agents. It converts any document — PDF, DOCX, slides, spreadsheets, images, audio, even whole websites — into clean, structured Markdown, then exposes it over MCP so your agent can search your FRDs, specs, and API docs on demand instead of re-reading (or forgetting) them.
Agent search: query-tailored web/news/paper/podcast segments, not full pages or links.
Reliable PDF table extraction. Pass a URL, get structured JSON tables with citations.
Connect Claude, Cursor, or ChatGPT to your business data. Ask questions, get answers.
ContextBook is an open-source MCP server that gives AI tools a persistent, searchable context library. Store information as Books and Pages, retrieve exactly what's needed via natural-language semantic search - injected on demand, not pre-loaded. Works with Cursor, Claude, Windsurf, and any MCP-compatible client. Self-hostable, MIT licensed.
Make your knowledge agent-ready. Connect docs from Confluence, Notion, GitHub, Dropbox, or Google Drive — any AI agent searches them via one MCP endpoint. 3 retrieval modes: vector search, broad search, and full document access. The agent decides how deep to dig.
Semantic memory for AI agents. Store, recall, and forget memories with entity extraction and fact-based retrieval. Self-hostable, privacy-first, works with any LLM.
Provides metadata information to AI agents through the search API.
The Needle MCP server enables semantic search on documents stored in files like PDFs, DOCX, and XLSX by connecting AI applications to external data sources. It provides capabilities to create and manage document collections, perform natural language searches on stored content, and retrieve relevant information without requiring exact keyword matches.