cocoindex MCP
Uses Hugging Face sentence-transformers models to generate embeddings for documents and queries, enabling similarity search.
Provides semantic search over indexed repositories and documents stored in a PostgreSQL database with the pgvector extension.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@cocoindex MCPsearch for error handling patterns in the codebase"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
cocoindex MCP
An MCP server that incrementally indexes repositories and documents into a Postgres + pgvector store using CocoIndex, and exposes semantic search over them.
The pipeline is source → extract (format registry) → chunk → embed → store:
Sources (
src/mcp_coco/sources.py) — local filesystem today, as two profiles:repo(code-aware, vendored dirs excluded) anddocument(markdown/text/pdf, prose chunking).Formats (
src/mcp_coco/formats.py) — a registry mapping a file to normalized text. PDF (viapymupdf) is just one handler; add a format by registering one.Indexer (
src/mcp_coco/indexer.py) — the CocoIndex app: chunk + embed (sentence-transformers) and declare rows into onedoc_embeddingstable.Search (
src/mcp_coco/db.py) — embeds the query and runs a pgvector similarity search.
CocoIndex tracks its incremental state in a local LMDB file (COCOINDEX_DB), so
re-indexing only reprocesses what changed and removes rows for deleted files.
Prerequisites
uv (Python package manager)
just (task runner, optional but convenient)
A Postgres instance with pgvector
Docker (if you want to run pgvector via the included compose file)
Related MCP server: ragi
Quick start (local)
1. Start a pgvector database
If you already have a Postgres instance with pgvector, skip this step and set
DATABASE_URL accordingly.
Otherwise, use the included compose file:
docker compose up -dThis starts pgvector on localhost:5432 with user/password/db all set to cocoindex.
2. Install dependencies
uv sync3. Configure
cp .env.example .envEdit .env and set DATABASE_URL to point at your Postgres instance. For the
Docker-based database:
DATABASE_URL=postgresql://cocoindex:cocoindex@localhost:5432/cocoindexOptional settings:
Variable | Default | Description |
|
| Embedding model for indexing and search |
|
| Cross-encoder model for result re-ranking |
|
| Postgres table name |
|
| Path to CocoIndex incremental state store |
4. Verify the database connection
just init5. Index something
just index ./path/to/repo repo
just index ./path/to/docs documentThe first run downloads the embedding model (~80 MB) from Hugging Face.
6. Search
just search "how does authentication work"Using with Coding Agents
Add the MCP server to your Claude Code settings
(~/.claude/settings.json for global, or .claude/settings.json in a project):
{
"mcpServers": {
"cocoindex": {
"command": "uv",
"args": ["run", "--directory", "/absolute/path/to/cocoindex-mcp", "mcp-coco-server"],
"env": {
"DATABASE_URL": "postgresql://cocoindex:cocoindex@localhost:5432/cocoindex",
"COCOINDEX_DB": "/absolute/path/to/cocoindex-mcp/.cocoindex/state.db"
}
}
}
}Replace /absolute/path/to/cocoindex-mcp with the actual path to this repository.
If your Postgres instance is elsewhere (e.g. a cloud-hosted database), adjust
DATABASE_URL accordingly. It is highly encouraged to pass your authentication information through env vars, do NOT hardcode into the connection string!
Once configured, Claude Code can use these tools:
Tool | Description |
| Index a code repository |
| Index a document collection |
| Semantic search — returns condensed summaries and a |
| Retrieve full details for specific results from a previous search |
Two-stage search
To keep context lean, search writes full results to a temporary JSON file
and returns only condensed summaries (~80-char excerpts) inline. The caller
triages from the summary, then uses read_search_results to fetch full
details for the results it actually needs.
By default, read_search_results re-ranks the selected results using a
cross-encoder model (cross-encoder/ms-marco-MiniLM-L-6-v2) for more
accurate relevance ordering. Disable with rerank=false. The model is
configurable via the RERANK_MODEL environment variable.
Development (devcontainer)
Open this folder in VS Code and Reopen in Container (Dev Containers). The
dbservice starts automatically alongside the app container.Run the preflight check:
just install just initCopy
.env.exampleto.envto customize settings. Inside the devcontainer the database hostname isdb(the default).
just recipes
just index <path> [repo|document|auto] # index a path
just index-repo <path> # index as code repository
just index-docs <path> # index as document collection
just search "query" [limit] # semantic search
just drop <path> [repo|document|auto] # remove a source from the index
just visualize_index # show a map of what's indexed
just serve # run the MCP server over stdio
just test # run tests
just lint # run ruffMaintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Tools
Related MCP Servers
- Flicense-qualityDmaintenanceLocal MCP server that provides semantic search (RAG) over code repositories, enabling AI clients like Claude and Gemini to access project context without manual re-upload.Last updated
- AlicenseAqualityBmaintenanceLocal-first RAG indexing and semantic search MCP server. Enables document retrieval and context-aware queries using local embedding models.Last updated325MIT
- Alicense-qualityDmaintenanceAn MCP server that indexes documents and serves relevant context to LLMs via Retrieval Augmented Generation (RAG).Last updated4336MIT
- Alicense-qualityFmaintenanceMCP server for semantic code search that indexes your codebase and allows AI editors to search using natural language queries.Last updated5854MIT
Related MCP Connectors
Local-first RAG engine with MCP server for AI agent integration.
Remote ChromaDB vector database MCP server with streamable HTTP transport
An MCP server that gives your AI access to the source code and docs of all public github repos
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/datarian/mcp-coco'
If you have feedback or need assistance with the MCP directory API, please join our Discord server