"Information about Word DOCX files" matching MCP connectors:
GET /v1/connectors – MCP directory API referenceMatching Connector Tools:
Unstructured document processing for LLM pipelines. Upload as PDF/DOCX/TXT any supported files, extract structured data (PII-redacted), build LLM-ready datasets, and search/export results — all via MCP tools (document.process, job.status, job.result, dataset.build, dataset.search, dataset.export).
PDF, Word, Excel, CSV, HTML and XML to clean Markdown for LLMs and RAG, with token counts.
One semantic search across your sites, Drive, Notion, email, files and Basecamp.
Your private knowledge base: upload documents (.md, .txt, .docx, PDF, images), the platform indexes
Manage your Mistral platform — models, files, batch jobs, agents and RAG document libraries.
Generate 18 AI readiness files (llms.txt, ai.txt, RAG indexes, schema) for any website.
Parse PDF/Word/PPT/HTML to Markdown; tables as JSON, image extraction, RAG chunking, page ranges.
PDF, Word, PowerPoint, Excel, HTML, EPUB to Markdown: OCR, page ranges, tables, RAG chunking
Rafter holds a team's durable knowledge — skills, agents and memory files, each versioned — and serves it to AI tools over MCP. Agents search across the team's artifacts before answering questions about how the team works or what was decided, fetch full artifact text along with its citation edges (cites, cited_by, links) to explore related material, and write new learnings back as memory. Also covers workspace, team and membership management.
LLMtoMD is the memory layer for AI coding agents. It converts any document — PDF, DOCX, slides, spreadsheets, images, audio, even whole websites — into clean, structured Markdown, then exposes it over MCP so your agent can search your FRDs, specs, and API docs on demand instead of re-reading (or forgetting) them.
ContextBook is an open-source MCP server that gives AI tools a persistent, searchable context library. Store information as Books and Pages, retrieve exactly what's needed via natural-language semantic search - injected on demand, not pre-loaded. Works with Cursor, Claude, Windsurf, and any MCP-compatible client. Self-hostable, MIT licensed.
MCP-native knowledge base for AI agents — vault-scoped docs, tables, and files, git-versioned, with hybrid search (BM25 + pgvector dense + reranker) and an event stream so external consolidators / gardeners stay decoupled.
Answers questions about a business using its indexed website content, with cited sources.
Vector RAG store for Word/Excel/PDF/PowerPoint. Break-even pricing, $5 per 5,700 pages.
The Needle MCP server enables semantic search on documents stored in files like PDFs, DOCX, and XLSX by connecting AI applications to external data sources. It provides capabilities to create and manage document collections, perform natural language searches on stored content, and retrieve relevant information without requiring exact keyword matches.
- cdgcsearchmetadataOAuth unavailableio.github.poojaBjAcharya
Provides metadata information to AI agents through the search API.