Skip to main content
Glama
cybernuki
by cybernuki

Context Engine MCP

A code-aware Context Engine accessible via the Model Context Protocol (MCP). Indexes repositories, maintains incremental sync, and returns highly relevant code snippets for any code generation or comprehension task.

Inspired by Augment Code's Context Engine. Built for any MCP client — Claude, Cursor, OpenCode, VS Code, etc.

Architecture at a glance

flowchart LR
    client["MCP client<br/>(Claude, Cursor, OpenCode, ...)"]

    subgraph server["context-engine MCP server"]
        direction TB
        transport["Transport<br/>stdio or Streamable HTTP + bearer auth"]
        tools["Tools<br/>index / search / get_file_context<br/>refresh / list / compare_branches"]
        security["Path security<br/>REPO_ROOTS allow-list"]
        transport --> security --> tools
    end

    subgraph pipeline["Indexing pipeline"]
        direction LR
        discover["Discover<br/>.gitignore aware"] --> hash["Hash + diff<br/>vs. saved state"]
        hash --> chunk["AST chunking<br/>tree-sitter"]
        chunk --> embed["Local embeddings<br/>fastembed / ONNX"]
    end

    qdrant[("Qdrant<br/>one collection per repo@branch")]

    client -- "MCP (stdio / HTTP)" --> transport
    tools --> pipeline
    embed -- "upsert vectors + payload" --> qdrant
    tools -- "embed query, similarity search" --> qdrant

Related MCP server: codesteer-atlas

Features

  • 6 MCP Tools: index_repository, search_codebase, get_file_context, refresh_changed_files, list_repositories, compare_branches

  • AST-aware code chunking — splits at function/class/method boundaries using tree-sitter via code-chunk

  • Local embeddings — no API keys needed, runs on CPU via fastembed (ONNX Runtime)

  • Qdrant BYOK — bring your own Qdrant instance (cloud or local)

  • Incremental indexing — Discover → Filter → Hash → Diff → Chunk → Embed → Upsert → Save

  • Hybrid search — semantic similarity + keyword filtering for precise symbol lookups

  • 60+ languages — TypeScript, JavaScript, Python, Go, Rust, Java, and many more

  • MCP Prompts — built-in guidance for LLMs on effective tool usage

Quick Start

Prerequisites

# Start a local Qdrant instance
docker run -p 6333:6333 -p 6334:6334 qdrant/qdrant

Install & Build

git clone https://github.com/cybernuki/context-engine-mcp.git
cd context-engine-mcp
npm install
npm run build

Configure your MCP Client

Add to your MCP client config (Claude Desktop, Cursor, OpenCode, etc.):

{
  "mcpServers": {
    "context-engine": {
      "command": "node",
      "args": ["/path/to/context-engine-mcp/dist/index.js"],
      "env": {
        "QDRANT_URL": "http://localhost:6333",
        "QDRANT_API_KEY": ""
      }
    }
  }
}

Tools

index_repository

Index or re-index a code repository.

{
  "repo_id": "my-app",
  "root_path": "/home/user/projects/my-app",
  "mode": "incremental"
}

search_codebase

Semantic search with optional filters and hybrid keyword matching.

{
  "repo_id": "my-app",
  "query": "authentication middleware",
  "top_k": 10,
  "score_threshold": 0.7,
  "filters": {
    "language": "typescript",
    "file_path_prefix": "src/auth/",
    "keyword": "verifyToken",
    "kind": "function"
  }
}

get_file_context

Read expanded file context after a search result.

{
  "repo_id": "my-app",
  "root_path": "/home/user/projects/my-app",
  "file_path": "src/auth/middleware.ts",
  "range": { "start_line": 1, "end_line": 50 }
}

refresh_changed_files

Quick reindex of files changed since last indexation.

{
  "repo_id": "my-app",
  "root_path": "/home/user/projects/my-app"
}

compare_branches

Compare a repo indexed per branch (repoIds <repo>@<branch>, e.g. my-app@main and my-app@development) by per-file file_hash.

Parameter

Description

repo

Base name, e.g. my-app

query

Optional. Semantic search in both branches; matching files are compared

file_paths

Optional. Explicit relative paths to compare

base_branch / head_branch

Optional. Default: from the repo registry (REPOS_REGISTRY), else main / development. A registry repo with head: null is compared as a single branch

limit

Search hits per branch (default 10)

Each file gets a status: same (in production), modified (integrated, pending release), only_head (new, pending release), only_base (removed in head), single_branch (repo has no integration branch; present in base/production). Unknown repos (with a registry) return an error listing the known ones. Hits from a query include lines and snippet. Uses Qdrant payload keys file_path and file_hash.

list_repositories

List all indexed repositories with stats, as JSON: indexed (collections and chunk counts) and registry (name, project, base, head, description, github from REPOS_REGISTRY, so clients can discover repos per project).

{}

Environment Variables

Variable

Required

Default

Description

QDRANT_URL

Yes

—

Qdrant cluster endpoint

QDRANT_API_KEY

No

—

Qdrant API key (if auth enabled)

EMBEDDING_MODEL

No

BAAI/bge-small-en-v1.5

Embedding model name

VECTOR_DIMENSION

No

384

Vector dimension (must match model)

DEFAULT_TOP_K

No

10

Default search result count

MAX_FILE_SIZE

No

1048576

Max file size to index (bytes)

REPOS_REGISTRY

No

—

Path to a repo registry JSON (deploy/repos.example.json shape: repos[] with name, project, branches.base, branches.head or null). Unset or missing file: behave as without a registry

INDEX_FILE_BATCH

No

25

Files per indexing group (chunk, embed/reuse, upsert, save state)

EMBED_BATCH_SIZE

No

32

Chunks per embedding call

EMBED_THREADS

No

1

ONNX intra-op threads for embedding

Indexing reuses vectors across branches: when a file has the same file_hash and chunk set in a sibling branch collection of the same repo, its points are copied instead of re-embedded (batched lookup per group; falls back to embedding on any error). Results report filesReused, chunksReused and chunksEmbedded.

HTTP Mode (remote / Docker)

Default transport is stdio. Set MCP_TRANSPORT=http to serve Streamable HTTP (stateless) at POST /mcp, plus GET /health (no auth).

Variable

Description

MCP_TRANSPORT

http to enable (default: stdio)

PORT

Listen port (default 8080)

MCP_AUTH_TOKEN

Required. Every /mcp request needs Authorization: Bearer <token> (constant-time compare). Server refuses to start if unset

REPO_ROOTS

JSON object repoId -> absolute path, e.g. {"my-app@main":"/repos/main","my-app@development":"/repos/dev"}

QDRANT_URL, QDRANT_API_KEY

Qdrant connection

EMBEDDING_CACHE_DIR

Fastembed model cache (Docker: /data/fastembed-cache, mount a volume)

ALLOW_REMOTE_FULL_INDEX

true to allow mode: "full" over HTTP (default: disabled)

In HTTP mode, root_path from clients is ignored: index_repository, refresh_changed_files and get_file_context resolve the root from REPO_ROOTS and reject unknown repo_ids. get_file_context rejects absolute paths, .. traversal (incl. percent-encoded) and symlink escapes. Full reindex (drops the collection) is disabled for remote callers by default; incremental stays available, and full rebuilds should run via the CLI/stdio on the server.

docker build -t context-engine .
docker run -p 8080:8080 -v ce-data:/data -v /srv/repos:/repos:ro   -e MCP_AUTH_TOKEN=... -e QDRANT_URL=... -e QDRANT_API_KEY=...   -e REPO_ROOTS='{"my-app@main":"/repos/main"}' context-engine

Indexer CLI

For systemd timers/cron. Runs the incremental pipeline once, logs JSON lines to stdout, exits non-zero on failure (needs QDRANT_URL, QDRANT_API_KEY):

node dist/cli.js refresh <repoId> <rootPath>
node dist/cli.js refresh my-app@main /srv/repos/main

Run npm test for the unit and property-based (fast-check) tests.

Source layout

src/
├── index.ts              # MCP server entry point (stdio or HTTP transport)
├── cli.ts                # One-shot indexer CLI (cron / systemd)
├── server.ts             # MCP server factory (tools + prompts)
├── config.ts             # BYOK environment configuration
├── registry.ts           # Optional repo registry (branch defaults)
├── prompts.ts            # MCP prompts for LLM guidance
├── http/                 # Streamable HTTP server + bearer-token auth
├── security/             # REPO_ROOTS parsing, path validation
├── compare/              # Per-file branch comparison
├── qdrant/
│   ├── client.ts         # Qdrant singleton, collection management
│   └── retry.ts          # Retry with backoff for transient errors
├── embeddings/
│   └── engine.ts         # fastembed wrapper (local ONNX embeddings)
├── chunking/
│   └── chunker.ts        # AST-aware code chunking (tree-sitter)
├── indexing/
│   ├── discovery.ts      # .gitignore-aware file walker
│   ├── state.ts          # Incremental state management
│   └── pipeline.ts       # Indexing pipeline orchestrator
└── tools/
    ├── errors.ts             # Shared error handling utilities
    ├── index-repository.ts
    ├── search-codebase.ts
    ├── get-file-context.ts
    ├── refresh-changed-files.ts
    ├── list-repositories.ts
    └── compare-branches.ts

How It Works

  1. Index: Walk the repo, parse code into AST chunks, embed locally, store in Qdrant

  2. Search: Embed the query, find similar vectors, apply filters, return formatted snippets

  3. Sync: Track file hashes, only re-process what changed (incremental indexing)

The indexing pipeline follows Augment's pattern: Discover → Filter → Hash → Diff → Chunk → Embed → Upsert → Save

Supported Embedding Models

Model

Dimensions

Speed

Quality

BAAI/bge-small-en-v1.5 (default)

384

Fast

Good

BAAI/bge-base-en-v1.5

768

Medium

Better

sentence-transformers/all-MiniLM-L6-v2

384

Fast

Good

intfloat/multilingual-e5-large

1024

Slow

Multilingual (no query/passage prefixes applied, see below)

When switching models, set EMBEDDING_MODEL and a matching VECTOR_DIMENSION, then run a full reindex (vectors from different models are not comparable).

Known limitation: non-English queries

The default bge-small-en-v1.5 is an English-only model. Queries written in other languages (for example Spanish or Portuguese) retrieve noticeably worse results, even when the code itself is in English. For multilingual teams use EMBEDDING_MODEL=intfloat/multilingual-e5-large with VECTOR_DIMENSION=1024; expect slower indexing and a much larger model download on CPU.

e5 prefixes are not implemented. The e5 model family is trained with query: / passage: input prefixes. This server embeds indexed chunks and search queries as raw text and adds no prefix, so retrieval with e5 works but is untested and likely below the model's best quality. Also, an unrecognized EMBEDDING_MODEL value silently falls back to bge-small-en-v1.5; only the models listed above are mapped.

Deploy on a VPS

The deploy/ folder is a complete, generic deployment for a small Linux box (2 vCPU / 4 GB is enough for a few mid-sized repos). It runs Qdrant and the HTTP server with docker compose, and a systemd timer that fetches your tracked branches every 10 minutes and reindexes only what changed.

# deploy/docker-compose.yml (abridged)
services:
  qdrant:
    image: qdrant/qdrant:v1.15.4
    environment:
      QDRANT__SERVICE__API_KEY: ${QDRANT__SERVICE__API_KEY}
    volumes: [qdrant-data:/qdrant/storage]
    networks: [internal]
  context-engine:
    build: ..
    env_file: .env
    environment: { REPOS_REGISTRY: /app/repos.json }
    volumes: [./repos:/repos:ro, ./repos.json:/app/repos.json:ro, ce-data:/data]
    networks: [internal]
  indexer:                # one-shot, bounded to 1 CPU / 1.5 GB
    image: context-engine:local
    profiles: [indexer]
    entrypoint: ["node", "dist/cli.js"]
    mem_limit: 1500m
    cpus: 1.0
# deploy/context-engine-refresh.timer
[Timer]
OnBootSec=5min
OnUnitActiveSec=10min
Persistent=true

Quick path:

cd deploy
cp ../.env.example .env && chmod 600 .env     # set MCP_AUTH_TOKEN and the Qdrant key
cp repos.example.json repos.json              # list your own repositories
./sync-repos.sh                               # deploy keys, clones, worktrees, REPO_ROOTS
docker compose up -d qdrant context-engine
./refresh.sh                                  # first index, then let the timer take over

See deploy/README.md for the full runbook (adding repos, SSH-tunnel bulk load, log rotation).

Security

  • Token auth on HTTP. MCP_TRANSPORT=http requires MCP_AUTH_TOKEN; the server refuses to start without it. Every /mcp request must carry Authorization: Bearer <token>, compared in constant time. Only /health is unauthenticated.

  • Server-side path resolution. In HTTP mode the client-supplied root_path is ignored. Roots come from the REPO_ROOTS allow-list and unknown repo_ids are rejected.

  • Path validation. get_file_context rejects absolute paths, .. traversal (including percent-encoded forms) and symlink escapes out of the repo root.

  • Destructive operations gated. A full reindex (which drops the collection) is disabled for remote callers unless ALLOW_REMOTE_FULL_INDEX=true.

  • Least privilege deployment. Repos are mounted read-only into the server, fetched with read-only deploy keys, and Qdrant is only reachable on the internal compose network.

Design decisions

  • Local embeddings. Code never leaves the machine and there is no per-token bill or API key to manage. bge-small-en-v1.5 runs on a plain CPU through ONNX Runtime and is small enough for a cheap VPS. The trade-off is quality and language coverage versus hosted models (see the known limitation above).

  • AST chunking. Fixed-size text windows cut functions in half and dilute the embedding. Splitting at function/class/method boundaries with tree-sitter keeps each chunk a coherent unit, so search hits map to something a model can actually read and edit.

  • Incremental hashing. Each file is content-hashed and compared with saved state, so a refresh only re-chunks and re-embeds what changed. State is written after every group of files, which makes a killed run resumable instead of starting over.

  • Branch-aware indexing. Each repo@branch is its own worktree and Qdrant collection. Files that are identical in a sibling branch (same path, hash and chunk set) have their vectors copied rather than re-embedded, so a second branch costs little CPU. The same per-file hashes power compare_branches, which tells you whether code is in production, only in the integration branch, or modified between them.

License

MIT

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    A high-performance MCP server for semantic search and codebase indexing using the Qdrant vector database. It features optimized embedding pipelines, AST-aware chunking, and git metadata enrichment for fast, privacy-focused local or remote search.
    1,052 npm
    23
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Local MCP server for semantic code search using Tree-sitter AST parsing, local embeddings, and hybrid search; enables indexing and querying codebases entirely offline.
    5
    MIT