Context Engine MCP
by cybernuki
README.md
# Context Engine MCP
A code-aware **Context Engine** accessible via the [Model Context Protocol (MCP)](https://modelcontextprotocol.io/). Indexes repositories, maintains incremental sync, and returns highly relevant code snippets for any code generation or comprehension task.
Inspired by [Augment Code's Context Engine](https://www.augmentcode.com/context-engine). Built for any MCP client — Claude, Cursor, OpenCode, VS Code, etc.
## Architecture at a glance
```mermaid
flowchart LR
client["MCP client<br/>(Claude, Cursor, OpenCode, ...)"]
subgraph server["context-engine MCP server"]
direction TB
transport["Transport<br/>stdio or Streamable HTTP + bearer auth"]
tools["Tools<br/>index / search / get_file_context<br/>refresh / list / compare_branches"]
security["Path security<br/>REPO_ROOTS allow-list"]
transport --> security --> tools
end
subgraph pipeline["Indexing pipeline"]
direction LR
discover["Discover<br/>.gitignore aware"] --> hash["Hash + diff<br/>vs. saved state"]
hash --> chunk["AST chunking<br/>tree-sitter"]
chunk --> embed["Local embeddings<br/>fastembed / ONNX"]
end
qdrant[("Qdrant<br/>one collection per repo@branch")]
client -- "MCP (stdio / HTTP)" --> transport
tools --> pipeline
embed -- "upsert vectors + payload" --> qdrant
tools -- "embed query, similarity search" --> qdrant
```
## Features
- **6 MCP Tools**: `index_repository`, `search_codebase`, `get_file_context`, `refresh_changed_files`, `list_repositories`, `compare_branches`
- **AST-aware code chunking** — splits at function/class/method boundaries using [tree-sitter](https://tree-sitter.github.io/) via [code-chunk](https://github.com/supermemoryai/code-chunk)
- **Local embeddings** — no API keys needed, runs on CPU via [fastembed](https://github.com/Anush008/fastembed-js) (ONNX Runtime)
- **Qdrant BYOK** — bring your own Qdrant instance (cloud or local)
- **Incremental indexing** — Discover → Filter → Hash → Diff → Chunk → Embed → Upsert → Save
- **Hybrid search** — semantic similarity + keyword filtering for precise symbol lookups
- **60+ languages** — TypeScript, JavaScript, Python, Go, Rust, Java, and many more
- **MCP Prompts** — built-in guidance for LLMs on effective tool usage
## Quick Start
### Prerequisites
- Node.js >= 18
- A Qdrant instance ([local Docker](https://qdrant.tech/documentation/quick-start/) or [Qdrant Cloud](https://cloud.qdrant.io/))
```bash
# Start a local Qdrant instance
docker run -p 6333:6333 -p 6334:6334 qdrant/qdrant
```
### Install & Build
```bash
git clone https://github.com/cybernuki/context-engine-mcp.git
cd context-engine-mcp
npm install
npm run build
```
### Configure your MCP Client
Add to your MCP client config (Claude Desktop, Cursor, OpenCode, etc.):
```jsonc
{
"mcpServers": {
"context-engine": {
"command": "node",
"args": ["/path/to/context-engine-mcp/dist/index.js"],
"env": {
"QDRANT_URL": "http://localhost:6333",
"QDRANT_API_KEY": ""
}
}
}
}
```
## Tools
### `index_repository`
Index or re-index a code repository.
```json
{
"repo_id": "my-app",
"root_path": "/home/user/projects/my-app",
"mode": "incremental"
}
```
### `search_codebase`
Semantic search with optional filters and hybrid keyword matching.
```json
{
"repo_id": "my-app",
"query": "authentication middleware",
"top_k": 10,
"score_threshold": 0.7,
"filters": {
"language": "typescript",
"file_path_prefix": "src/auth/",
"keyword": "verifyToken",
"kind": "function"
}
}
```
### `get_file_context`
Read expanded file context after a search result.
```json
{
"repo_id": "my-app",
"root_path": "/home/user/projects/my-app",
"file_path": "src/auth/middleware.ts",
"range": { "start_line": 1, "end_line": 50 }
}
```
### `refresh_changed_files`
Quick reindex of files changed since last indexation.
```json
{
"repo_id": "my-app",
"root_path": "/home/user/projects/my-app"
}
```
### `compare_branches`
Compare a repo indexed per branch (repoIds `<repo>@<branch>`, e.g. `my-app@main` and `my-app@development`) by per-file `file_hash`.
| Parameter | Description |
|-----------|-------------|
| `repo` | Base name, e.g. `my-app` |
| `query` | Optional. Semantic search in both branches; matching files are compared |
| `file_paths` | Optional. Explicit relative paths to compare |
| `base_branch` / `head_branch` | Optional. Default: from the repo registry (`REPOS_REGISTRY`), else `main` / `development`. A registry repo with `head: null` is compared as a single branch |
| `limit` | Search hits per branch (default 10) |
Each file gets a status: `same` (in production), `modified` (integrated, pending release), `only_head` (new, pending release), `only_base` (removed in head), `single_branch` (repo has no integration branch; present in base/production). Unknown repos (with a registry) return an error listing the known ones. Hits from a query include lines and snippet. Uses Qdrant payload keys `file_path` and `file_hash`.
### `list_repositories`
List all indexed repositories with stats, as JSON: `indexed` (collections and chunk counts) and `registry` (`name`, `project`, `base`, `head`, `description`, `github` from `REPOS_REGISTRY`, so clients can discover repos per project).
```json
{}
```
## Environment Variables
| Variable | Required | Default | Description |
|----------|----------|---------|-------------|
| `QDRANT_URL` | Yes | — | Qdrant cluster endpoint |
| `QDRANT_API_KEY` | No | — | Qdrant API key (if auth enabled) |
| `EMBEDDING_MODEL` | No | `BAAI/bge-small-en-v1.5` | Embedding model name |
| `VECTOR_DIMENSION` | No | `384` | Vector dimension (must match model) |
| `DEFAULT_TOP_K` | No | `10` | Default search result count |
| `MAX_FILE_SIZE` | No | `1048576` | Max file size to index (bytes) |
| `REPOS_REGISTRY` | No | — | Path to a repo registry JSON (`deploy/repos.example.json` shape: `repos[]` with `name`, `project`, `branches.base`, `branches.head` or null). Unset or missing file: behave as without a registry |
| `INDEX_FILE_BATCH` | No | `25` | Files per indexing group (chunk, embed/reuse, upsert, save state) |
| `EMBED_BATCH_SIZE` | No | `32` | Chunks per embedding call |
| `EMBED_THREADS` | No | `1` | ONNX intra-op threads for embedding |
Indexing reuses vectors across branches: when a file has the same `file_hash` and chunk set in a sibling branch collection of the same repo, its points are copied instead of re-embedded (batched lookup per group; falls back to embedding on any error). Results report `filesReused`, `chunksReused` and `chunksEmbedded`.
## HTTP Mode (remote / Docker)
Default transport is stdio. Set `MCP_TRANSPORT=http` to serve Streamable HTTP (stateless) at `POST /mcp`, plus `GET /health` (no auth).
| Variable | Description |
|----------|-------------|
| `MCP_TRANSPORT` | `http` to enable (default: stdio) |
| `PORT` | Listen port (default `8080`) |
| `MCP_AUTH_TOKEN` | Required. Every `/mcp` request needs `Authorization: Bearer <token>` (constant-time compare). Server refuses to start if unset |
| `REPO_ROOTS` | JSON object `repoId -> absolute path`, e.g. `{"my-app@main":"/repos/main","my-app@development":"/repos/dev"}` |
| `QDRANT_URL`, `QDRANT_API_KEY` | Qdrant connection |
| `EMBEDDING_CACHE_DIR` | Fastembed model cache (Docker: `/data/fastembed-cache`, mount a volume) |
| `ALLOW_REMOTE_FULL_INDEX` | `true` to allow `mode: "full"` over HTTP (default: disabled) |
In HTTP mode, `root_path` from clients is ignored: `index_repository`, `refresh_changed_files` and `get_file_context` resolve the root from `REPO_ROOTS` and reject unknown `repo_id`s. `get_file_context` rejects absolute paths, `..` traversal (incl. percent-encoded) and symlink escapes. Full reindex (drops the collection) is disabled for remote callers by default; incremental stays available, and full rebuilds should run via the CLI/stdio on the server.
```bash
docker build -t context-engine .
docker run -p 8080:8080 -v ce-data:/data -v /srv/repos:/repos:ro -e MCP_AUTH_TOKEN=... -e QDRANT_URL=... -e QDRANT_API_KEY=... -e REPO_ROOTS='{"my-app@main":"/repos/main"}' context-engine
```
## Indexer CLI
For systemd timers/cron. Runs the incremental pipeline once, logs JSON lines to stdout, exits non-zero on failure (needs `QDRANT_URL`, `QDRANT_API_KEY`):
```bash
node dist/cli.js refresh <repoId> <rootPath>
node dist/cli.js refresh my-app@main /srv/repos/main
```
Run `npm test` for the unit and property-based (fast-check) tests.
## Source layout
```
src/
├── index.ts # MCP server entry point (stdio or HTTP transport)
├── cli.ts # One-shot indexer CLI (cron / systemd)
├── server.ts # MCP server factory (tools + prompts)
├── config.ts # BYOK environment configuration
├── registry.ts # Optional repo registry (branch defaults)
├── prompts.ts # MCP prompts for LLM guidance
├── http/ # Streamable HTTP server + bearer-token auth
├── security/ # REPO_ROOTS parsing, path validation
├── compare/ # Per-file branch comparison
├── qdrant/
│ ├── client.ts # Qdrant singleton, collection management
│ └── retry.ts # Retry with backoff for transient errors
├── embeddings/
│ └── engine.ts # fastembed wrapper (local ONNX embeddings)
├── chunking/
│ └── chunker.ts # AST-aware code chunking (tree-sitter)
├── indexing/
│ ├── discovery.ts # .gitignore-aware file walker
│ ├── state.ts # Incremental state management
│ └── pipeline.ts # Indexing pipeline orchestrator
└── tools/
├── errors.ts # Shared error handling utilities
├── index-repository.ts
├── search-codebase.ts
├── get-file-context.ts
├── refresh-changed-files.ts
├── list-repositories.ts
└── compare-branches.ts
```
## How It Works
1. **Index**: Walk the repo, parse code into AST chunks, embed locally, store in Qdrant
2. **Search**: Embed the query, find similar vectors, apply filters, return formatted snippets
3. **Sync**: Track file hashes, only re-process what changed (incremental indexing)
The indexing pipeline follows Augment's pattern: **Discover → Filter → Hash → Diff → Chunk → Embed → Upsert → Save**
## Supported Embedding Models
| Model | Dimensions | Speed | Quality |
|-------|-----------|-------|---------|
| `BAAI/bge-small-en-v1.5` (default) | 384 | Fast | Good |
| `BAAI/bge-base-en-v1.5` | 768 | Medium | Better |
| `sentence-transformers/all-MiniLM-L6-v2` | 384 | Fast | Good |
| `intfloat/multilingual-e5-large` | 1024 | Slow | Multilingual (no query/passage prefixes applied, see below) |
When switching models, set `EMBEDDING_MODEL` and a matching `VECTOR_DIMENSION`, then run a full reindex (vectors from different models are not comparable).
### Known limitation: non-English queries
The default `bge-small-en-v1.5` is an English-only model. Queries written in other languages (for example Spanish or Portuguese) retrieve noticeably worse results, even when the code itself is in English. For multilingual teams use `EMBEDDING_MODEL=intfloat/multilingual-e5-large` with `VECTOR_DIMENSION=1024`; expect slower indexing and a much larger model download on CPU.
> **e5 prefixes are not implemented.** The e5 model family is trained with `query: ` / `passage: ` input prefixes. This server embeds indexed chunks and search queries as raw text and adds no prefix, so retrieval with e5 works but is untested and likely below the model's best quality. Also, an unrecognized `EMBEDDING_MODEL` value silently falls back to `bge-small-en-v1.5`; only the models listed above are mapped.
## Deploy on a VPS
The `deploy/` folder is a complete, generic deployment for a small Linux box (2 vCPU / 4 GB is enough for a few mid-sized repos). It runs Qdrant and the HTTP server with `docker compose`, and a systemd timer that fetches your tracked branches every 10 minutes and reindexes only what changed.
```yaml
# deploy/docker-compose.yml (abridged)
services:
qdrant:
image: qdrant/qdrant:v1.15.4
environment:
QDRANT__SERVICE__API_KEY: ${QDRANT__SERVICE__API_KEY}
volumes: [qdrant-data:/qdrant/storage]
networks: [internal]
context-engine:
build: ..
env_file: .env
environment: { REPOS_REGISTRY: /app/repos.json }
volumes: [./repos:/repos:ro, ./repos.json:/app/repos.json:ro, ce-data:/data]
networks: [internal]
indexer: # one-shot, bounded to 1 CPU / 1.5 GB
image: context-engine:local
profiles: [indexer]
entrypoint: ["node", "dist/cli.js"]
mem_limit: 1500m
cpus: 1.0
```
```ini
# deploy/context-engine-refresh.timer
[Timer]
OnBootSec=5min
OnUnitActiveSec=10min
Persistent=true
```
Quick path:
```bash
cd deploy
cp ../.env.example .env && chmod 600 .env # set MCP_AUTH_TOKEN and the Qdrant key
cp repos.example.json repos.json # list your own repositories
./sync-repos.sh # deploy keys, clones, worktrees, REPO_ROOTS
docker compose up -d qdrant context-engine
./refresh.sh # first index, then let the timer take over
```
See [`deploy/README.md`](deploy/README.md) for the full runbook (adding repos, SSH-tunnel bulk load, log rotation).
## Security
- **Token auth on HTTP.** `MCP_TRANSPORT=http` requires `MCP_AUTH_TOKEN`; the server refuses to start without it. Every `/mcp` request must carry `Authorization: Bearer <token>`, compared in constant time. Only `/health` is unauthenticated.
- **Server-side path resolution.** In HTTP mode the client-supplied `root_path` is ignored. Roots come from the `REPO_ROOTS` allow-list and unknown `repo_id`s are rejected.
- **Path validation.** `get_file_context` rejects absolute paths, `..` traversal (including percent-encoded forms) and symlink escapes out of the repo root.
- **Destructive operations gated.** A full reindex (which drops the collection) is disabled for remote callers unless `ALLOW_REMOTE_FULL_INDEX=true`.
- **Least privilege deployment.** Repos are mounted read-only into the server, fetched with read-only deploy keys, and Qdrant is only reachable on the internal compose network.
## Design decisions
- **Local embeddings.** Code never leaves the machine and there is no per-token bill or API key to manage. `bge-small-en-v1.5` runs on a plain CPU through ONNX Runtime and is small enough for a cheap VPS. The trade-off is quality and language coverage versus hosted models (see the known limitation above).
- **AST chunking.** Fixed-size text windows cut functions in half and dilute the embedding. Splitting at function/class/method boundaries with tree-sitter keeps each chunk a coherent unit, so search hits map to something a model can actually read and edit.
- **Incremental hashing.** Each file is content-hashed and compared with saved state, so a refresh only re-chunks and re-embeds what changed. State is written after every group of files, which makes a killed run resumable instead of starting over.
- **Branch-aware indexing.** Each `repo@branch` is its own worktree and Qdrant collection. Files that are identical in a sibling branch (same path, hash and chunk set) have their vectors copied rather than re-embedded, so a second branch costs little CPU. The same per-file hashes power `compare_branches`, which tells you whether code is in production, only in the integration branch, or modified between them.
## License
MIT
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues