CorpusGate
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@CorpusGatesearch my local docs for the Q3 financial report highlights"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
CorpusGate
An LLM-ready private document gateway powered by MarkItDown and MCP.
Convert, index and retrieve private documents for AI tools without sending document content to third-party services.
CorpusGate is a general-purpose, self-hosted Document MCP Server for individuals and teams that need controlled AI-tool access to documents on their own infrastructure. MarkItDown converts supported files into reusable Markdown. The gateway then chunks and indexes that Markdown, and MCP returns only the relevant, source-attributed chunks under server-enforced budgets.
Self-hosting keeps source documents, generated Markdown, queries, metadata, and indexes under the operator's control. Token reduction comes from bounded retrieval and chunk selection—not from MarkItDown alone. This project is not a chatbot, an LLM answer generator, a contract-analysis product, a SaaS platform, or a user-facing document panel.
Features
PDF, DOCX, PPTX, XLSX, TXT, Markdown, and HTML conversion through Microsoft MarkItDown.
Persistent Markdown cache, token-aware heading-preserving chunks, and SHA-256 deduplication.
SQLite FTS5/BM25 lexical search in the lightweight default installation.
Optional CPU-only multilingual semantic search and RRF hybrid retrieval using local embeddings.
Bounded MCP responses with source/position metadata, cursors, deduplication, and neighbor limits.
REST API-key and MCP Bearer authentication, safe UUID storage, path/symlink protections, rate limits, and structured content-safe logs.
Hardened Docker Compose deployment for AMD64/ARM64, Oracle Cloud, Tailscale, or Caddy HTTPS.
Setup helper, operational doctor/scan/reindex/backup commands, versioned SQLite schema, and CI.
Related MCP server: rag-retriever-mcp
How it works
REST upload or read-only inbox scan
│
├─ type, signature, size, path, and free-space validation
├─ UUID storage + SHA-256 ── unchanged? ── reuse cached/indexed record
│
└─ MarkItDown ──> persistent Markdown ──> token-aware chunks
│
┌────────────────────┴────────────────────┐
│ │
SQLite FTS5 / BM25 optional local embeddings
│ + private Qdrant
└────────────────────┬────────────────────┘
│
ranking → dedup → token/char budget → MCPThe code keeps parser, storage, repository, chunking, embedding, vector-store, and retrieval contracts behind interfaces without turning the single-server product into a distributed system. Documents remain the primary data source; Markdown and vector indexes can be rebuilt.
Supported formats
Format | Extensions | Notes |
| Text-based PDFs; no external OCR in | |
Word |
| Office archive structure is checked. |
PowerPoint |
| Slide markers are preserved when MarkItDown emits them. |
Excel |
| Sheet headings are carried into chunk metadata when available. |
Text |
| UTF-8. |
Markdown |
| UTF-8 and heading aware. |
HTML |
| UTF-8; fetching remote URLs is intentionally unsupported. |
Encrypted, corrupted, scanned-image-only, or converter-unsupported files fail safely without stopping other documents.
Quick start
Requirements: Docker Engine with Compose v2 and OpenSSL. Host Python is not required.
git clone https://github.com/mustafa0zdemir/corpusgate.git
cd corpusgate
./corpusgate init
./corpusgate up
curl --fail http://127.0.0.1:8000/health
./corpusgate doctor./corpusgate init creates persistent/inbox folders, copies .env.example only when .env does not
exist, generates separate random REST/MCP credentials without printing them, checks Docker/Compose
and the selected port, and validates Compose. It never overwrites an existing .env.
The equivalent manual flow is to copy .env.example to .env, replace both credential
placeholders with different openssl rand -hex 32 values, create documents/, and run
docker compose up -d. Never commit .env.
Optional local semantic/hybrid retrieval is also one operation after initialization:
./corpusgate init --semantic
./corpusgate up --semanticThe first semantic start downloads the model into a persistent cache and then starts the gateway offline with Qdrant on the internal Docker network. Later starts reuse both model and vector volumes. The lexical installation does not install or run either semantic component.
Add documents
The simplest operator workflow uses the read-only host inbox:
cp examples/documents/* documents/
./corpusgate scan
./corpusgate list-documentsThe scan skips hidden/system/temporary files, unsupported types, directories, and symlinks. Input
files remain in documents/; private UUID copies are stored in the persistent source volume.
For a complete lexical→semantic/hybrid→MCP walkthrough, use the
synthetic demo.
For a one-file workflow that AI tools can invoke without putting file bytes into the model context,
stream the local file directly to the running REST API (requires curl):
./corpusgate upload /absolute/path/to/document.pdfThe command reads the REST key from CORPUSGATE_CLIENT_API_KEY, CORPUSGATE_API_KEY, or the local
.env, never prints it, rejects redirects and insecure remote HTTP, and returns only the API's
upload metadata. For a remote private server, pass
--url https://YOUR-NODE.YOUR-TAILNET.ts.net.
REST upload is available for applications:
export CORPUSGATE_CLIENT_API_KEY='value-from-your-env'
curl --fail -X POST http://127.0.0.1:8000/api/v1/documents \
-H "X-API-Key: ${CORPUSGATE_CLIENT_API_KEY}" \
-F 'file=@examples/documents/private-network-guide.md'REST also provides paginated metadata, Markdown, chunks, lexical search, and deletion under
/api/v1/documents. Interactive OpenAPI documentation is at /docs; protected operations still
require X-API-Key.
Connect an MCP client
The remote endpoint is https://YOUR_PRIVATE_OR_PUBLIC_HOST/mcp and every MCP request needs:
Authorization: Bearer YOUR_MCP_TOKENUse Tailscale Serve as the recommended private route. Caddy HTTPS is the public alternative; the gateway port remains bound to host loopback in both cases. See the verified field mapping, Inspector command, Tailscale/HTTPS examples, and troubleshooting in the MCP connection guide. Do not copy an unverified client-specific JSON wrapper or save a token in source control.
MCP tools
Tool | Purpose | Limits and behavior |
| Discover metadata without text. |
|
| Inspect one source/status/cache record. | Returns no document content. |
| Search when the source document is unknown. | Mode/filters/top-k/budgets/cursor. |
| Search one known document. | Optional limited neighbors. |
| Build a small context set from an allowlist. | Deduplicated and budgeted. |
| Read consecutive chunks after locating a position. | Chunk cursor and hard budgets; never raw file. |
| Idempotently repair one stored document's lexical/optional vector index. | Returns maintenance counts, no content; does not upload/delete/reconvert. |
Retrieval items consistently include document_id, document_name, chunk_id, heading,
position, relevance/rank fields, bounded content, content_length, and retrieval-mode
metadata. Empty searches return an empty items list, applied budgets, metrics, and no cursor.
Invalid modes, cursors, filters, document IDs, or over-limit values produce controlled tool errors.
Upload and delete remain REST-only.
Recommended flow:
AI tool → search_document(query, top_k=3, max_tokens=600)
→ ranked chunks + source positions + actual retrieval mode
→ optional bounded get_document_sectionLexical, semantic, and hybrid search
lexicalis the production default: SQLite FTS5 with heading-weighted BM25 preserves exact identifiers and phrases without another service.semanticembeds queries/chunks locally with the configurable multilingual CPU model and stores vectors in private Qdrant.hybridmerges independent lexical and semantic ranks with Reciprocal Rank Fusion; duplicate chunks are returned once and exact lexical matches are not discarded.lexical_fallbackis reported when semantic/hybrid was requested but the optional local model, vector store, or index is unavailable and fallback is enabled.
The default model is Apache-2.0-licensed
sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2,
a 384-dimensional multilingual model used through FastEmbed/ONNX on CPU. See model replacement,
offline transfer, reindex rules, measurements, and Oracle memory guidance in
semantic search.
Token optimization
MarkItDown makes diverse files consistently parseable; it does not by itself guarantee fewer
tokens. The gateway reduces returned context by caching conversion once, ranking chunks, omitting
duplicates, enforcing top_k, max_chars, and estimated max_tokens, limiting neighbors, and
paginating long result sets. It never exposes a default MCP tool that returns a complete document
or raw file.
Token counts are a local deterministic estimate, not a provider-specific billing tokenizer. The repeatable synthetic measurement and its exact scope are in the retrieval report; no universal saving percentage is claimed.
Security and privacy
No telemetry, document text, query text, cloud embedding API, or required LLM provider.
Source files use UUID paths; filename traversal, absolute paths, symlink escapes, hidden/temp files, MIME/signature mismatches, archive expansion, upload size, and low disk are checked.
REST uses API keys; remote MCP uses constant-time compared Bearer tokens from environment or a Docker secret. Multiple current/previous tokens allow rotation.
Structured logs contain allowlisted operation metadata, never document content, credentials, complete queries, or client-visible stack traces.
The gateway container is non-root, capability-free,
no-new-privileges, read-only-root, and has explicit writable volumes/tmpfs plus resource/log limits.Base Compose publishes only
127.0.0.1:8000; Qdrant is internal-only. Public deployment requires Caddy TLS and keeps Bearer authentication/rate/response limits.
Stored data comprises private UUID source copies, generated Markdown, SQLite metadata/chunks/FTS, optional local vectors/model cache, backups, and operator configuration. To erase data, delete documents through REST first; remove persistent volumes only after an explicit backup and shutdown. See SECURITY.md for private vulnerability reporting.
Oracle Cloud deployment
The recommended Oracle Ubuntu deployment binds the app to loopback and uses Tailscale Serve for
tailnet-only HTTPS. A Caddy public profile is documented for cases that require a domain. Oracle
Security Lists/NSGs must never expose TCP 8000 or Qdrant 6333.
VM preparation, AMD64/Ampere ARM64 notes, Docker installation, filesystem ownership, secrets, firewalls, Tailscale/Caddy, logging, update, backup, restore, and troubleshooting are in the Oracle deployment guide.
Configuration
All application environment variables, defaults, requirements, ranges, examples, and security
effects are listed in the configuration contract and .env.example.
Startup rejects missing/short credentials, invalid ports/paths, impossible chunk/budget
relationships, unsupported retrieval modes, and invalid semantic vector-store configuration
without echoing secret values.
Operational commands:
./corpusgate version
./corpusgate status
./corpusgate doctor
./corpusgate mcp-smoke
./corpusgate upload /absolute/path/to/document.pdf
./corpusgate scan
./corpusgate reindex
./corpusgate reindex --semantic
./corpusgate list-documents --limit 20 --offset 0Backup and restore
./corpusgate backup
./corpusgate restore /backups/corpusgate-backup-TIMESTAMP.tar.gz --confirm-restoreRestore replaces current persistent data and therefore requires an explicit confirmation flag and
a stopped writer in production. Backups include private source storage, Markdown cache, a
transactionally copied SQLite database, manifest, and secret-free config example. Keep .env and
token files in a separate encrypted secret backup. Vector data is rebuildable from chunks.
Update and rollback
Learn the current version, backup, select the reviewed tag/image, run the versioned idempotent migration, restart, check ready/MCP, and retain the backup until validation completes. A database created by a newer incompatible application is rejected instead of silently modified.
Exact commands and the safe rollback/restore path are in
update and rollback. Never run docker compose down -v during a
normal update.
The repository also contains a manual, approval-gated GHCR workflow. Stable tag, moving minor, and
latest behavior are defined in the container publishing policy;
no image has been published by this sprint.
Troubleshooting
./corpusgate doctor: validates config, storage permissions, SQLite/schema, disk, optional model/vector state, service readiness, and version without dumping secrets.401: use RESTX-API-Keyor MCPAuthorization: Bearer, not the other credential type.Host rejection: add the exact Tailscale/domain host to
CORPUSGATE_ALLOWED_HOSTSand recreate gateway.507: free disk or review the reserved disk threshold before retrying ingestion.lexical_fallback: inspect model cache and Qdrant health; lexical retrieval remains available.Conversion failure: confirm supported extension, MIME/signature, UTF-8/Office archive integrity, size, encryption, and whether the PDF contains text.
Logs:
./corpusgate logs --tail=100; sanitize output before sharing.
See SUPPORT.md and the deployment-specific troubleshooting guide before opening an issue.
Compatibility
Environment | v0.1.0 status |
Python | Runtime image uses Python 3.12; automated tests target 3.12. |
| Runtime and semantic image build/run validated on an ARM64 Docker host. |
| Multi-architecture Buildx CI target; release requires checklist validation. |
Oracle Cloud Ubuntu | Deployment contract targets Ubuntu 24.04/Ampere; fresh VM validation remains a release checklist item. |
Docker / Compose | ARM64 flow tested with Engine 29.6.2 and Compose 5.3.1; Compose v2 is required. |
Lexical search | Default image; no semantic service required. |
Semantic search | Optional image/Qdrant/model volumes; CPU-only tested on ARM64. |
Offline mode | Lexical is offline; semantic is offline after the one-time model cache fill. |
Untested platforms are not presented as supported. Review the release checklist before publishing artifacts.
Limitations
Single-node SQLite is not a high-availability or multi-writer database.
Upload conversion is synchronous in
0.1.0; large documents may need longer client/proxy timeouts.No OCR, cloud storage adapter, user accounts, UI, answer generation, reranker, fine-tuning, or SaaS control plane.
Approximate token budgets may differ from a specific LLM tokenizer.
Semantic model download needs temporary outbound access unless the cache is transferred offline.
Roadmap
Background conversion jobs without making Redis mandatory for single-node users.
Optional PostgreSQL/pgvector and object-storage adapters behind existing interfaces.
More converter metadata extraction and operator-controlled OCR adapter.
Signed releases, SBOM/provenance, expanded cross-architecture and upgrade fixtures.
Contributing
Read CONTRIBUTING.md, follow CODE_OF_CONDUCT.md, add tests, and use only synthetic non-sensitive fixtures. Security reports must use the private route in SECURITY.md, never a public issue.
License
CorpusGate is available under the existing MIT License. Third-party libraries and the optional embedding model retain their own licenses.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- FlicenseNot gradedqualityBmaintenanceEnables any MCP-compatible AI assistant to search, filter, and retrieve information from a local document collection using a hybrid search pipeline with vector, BM25, reranking, and LLM enrichment.4
- FlicenseAqualityBmaintenanceA local-first document retrieval engine that mounts as an MCP tool for agents to index files, search for relevant passages, and let the agent's own LLM answer.4
- FlicenseNot gradedqualityCmaintenanceEnables users to build and query a private knowledge base by uploading documents, which are embedded and stored locally, then accessible via MCP for semantic search and retrieval.
- FlicenseNot gradedqualityCmaintenanceEnables local document question-answering and retrieval via MCP, supporting multi-turn conversation, intent recognition, and tools for document search, Q&A, and summarization.5
Related MCP Connectors
Multi-engine search for AI agents. Trust scoring, local corpus, MCP-native. Self-hostable, BYOK.
Search your AI chat history (ChatGPT, Claude, Codex) from any MCP client. Remote, private, read-only
Agentic search over your Dewey document collections from any MCP-compatible client.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/mustafa0zdemir/corpusgate'
If you have feedback or need assistance with the MCP directory API, please join our Discord server