ContextTree MCP
<div align="center">
# π³ ContextTree MCP
**Local Deep Semantic & Hybrid Code Search Engine for AI Assistants**
*Powered by tree-sitter AST logical parsing, local embeddings, 3-layer RRF ranking & Cross-Encoder re-ranking.*
[](https://github.com/chelslava/mcp-context-tree/releases)
[](https://www.python.org/)
[](https://modelcontextprotocol.io)
[](https://github.com/chelslava/mcp-context-tree/actions)
[](https://github.com/astral-sh/ruff)
[](LICENSE)
[](#privacy--security)
[](README.ru.md)
<p align="center">
<a href="#key-features">Key Features</a> β’
<a href="#supported-languages">Languages</a> β’
<a href="#architecture">Architecture</a> β’
<a href="#quick-start">Quick Start</a> β’
<a href="#client-configuration">Client Configs</a> β’
<a href="#mcp-tools-reference">MCP Tools</a> β’
<a href="#search-algorithm">Search Algorithm</a>
</p>
</div>
---
## π‘ Why ContextTree MCP?
Standard semantic search tools split code into arbitrary line or token windows, breaking function contexts and hallucinating definitions. **ContextTree MCP** provides LLMs with true **structural intelligence** of your codebase:
- π§© **AST Logical Block Extraction:** Indexes complete, meaningful units (functions, methods, classes, structs, traits) preserving docstrings and signatures.
- π’ **Multi-Repository Unified Indexing:** Seamlessly index and search across multiple repositories or multi-root workspaces in a single session.
- ποΈ **Embedding Quantization (INT8 & Binary):** 4x to 32x RAM and storage compression with scalar/binary quantized vector representations.
- π **Language Server Protocol (LSP) Bridge:** Compiler-grade symbol definitions, hover documentation, and references across local language servers.
- β‘ **3-Layer Hybrid Search (RRF):** Blends dense vectors (`sentence-transformers/all-MiniLM-L6-v2`), BM25 lexical token matching (`camelCase`/`snake_case`), and **Call-Graph In-Degree Ranking**.
- π― **Cross-Encoder 2nd-Stage Re-ranking:** Joint cross-attention re-scoring (`rerank=True`) for maximum precision on nuanced queries.
- π **Zero-Hallucination Code Navigation:** Real AST call-site tracking (`find_ast_usages`) and cross-file definition jump (`go_to_definition`).
- π **Incremental Indexing & Watch Mode:** SHA-256 state tracking with 500ms debounced filesystem watcher across multiple workspaces.
- π **Flexible Transports:** Standard `stdio`, `SSE` over HTTP, and `Streamable HTTP`.
- π **100% Offline & Private:** Zero cloud dependencies, zero external API calls, zero telemetry.
---
## π Supported Languages (12 Languages)
| Language | Extensions | Extracted AST Constructs |
|:---|:---|:---|
| **Python** | `.py` | Functions, decorated definitions, classes, methods, docstrings (PEP-257) |
| **TypeScript / TSX** | `.ts`, `.tsx`, `.mts`, `.cts` | Functions, arrow functions, methods, class/interface signatures, JSDoc |
| **JavaScript / JSX** | `.js`, `.jsx`, `.mjs`, `.cjs` | Functions, arrow functions, methods, class signatures, JSDoc |
| **Go** | `.go` | Functions, receiver methods, struct/interface types, package comments |
| **Rust** | `.rs` | Functions, `impl` methods, structs, traits, `///` documentation |
| **C#** | `.cs` | Methods, constructors, classes, interfaces, structs, `/// <summary>` XML-docs |
| **Java** | `.java` | Methods, constructors, classes, interfaces, records, Javadoc |
| **C** | `.c`, `.h` | Functions, structs, unions, enums, declarator unpacking, comments |
| **C++** | `.cpp`, `.hpp`, `.cc`, `.cxx`, `.hh`, `.hxx` | Methods, classes, structs, namespaces, destructors, doc comments |
| **Kotlin** | `.kt`, `.kts` | Functions, classes, objects, member methods, KDoc comments |
| **Swift** | `.swift` | Functions, methods, classes, structs, protocols, enums, Swift-doc |
---
## ποΈ Architecture
```mermaid
flowchart TB
subgraph Client["π€ AI Assistant Client"]
Claude["Claude Desktop / Cursor / Antigravity / OpenCode"]
end
subgraph Server["π³ ContextTree MCP Server"]
Transport["Transport Layer (Stdio / SSE / HTTP)"]
Tools["MCP Tools (search, usages, definition, index)"]
subgraph Pipeline["Indexing & Search Pipeline"]
TreeSitter["Tree-sitter AST Parser (12 Grammars)"]
Chunker["Logical Block Chunker (Signatures + Docs)"]
BM25["In-Memory BM25 Index (Cached)"]
VectorStore["ChromaDB Vector Store (384d Embeddings)"]
CallGraph["Call-Graph In-Degree Frequency"]
RRF["3-Layer Reciprocal Rank Fusion"]
CrossEncoder["Cross-Encoder Re-ranker (ms-marco-MiniLM)"]
end
end
subgraph Workspace["π» Local Workspace Files"]
SourceFiles["Source Code (.py, .ts, .go, .rs, .cpp, .kt, ...)"]
State[".chroma/index_state.json (SHA-256 Fast Path)"]
end
Claude <--> Transport
Transport <--> Tools
Tools <--> Pipeline
Pipeline <--> Workspace
```
---
## π Quick Start
### Prerequisites
- Python **3.12+**
- [`uv`](https://docs.astral.sh/uv/) (strongly recommended) or standard `pip`
### 1. Installation
```bash
# Clone the repository
git clone https://github.com/chelslava/mcp-context-tree.git
cd mcp-context-tree
# Install dependencies and local package via uv
uv sync
```
### 2. Running ContextTree MCP
```bash
# Standard MCP stdio mode (default for AI desktop clients)
uv run context-tree
# Server-Sent Events (SSE) HTTP transport on port 8000
uv run context-tree --transport sse --host 127.0.0.1 --port 8000
# Streamable HTTP transport
uv run context-tree --transport streamable-http --port 8000
# Standalone Watch Mode (continuously indexes workspace on save)
uv run context-tree --watch /path/to/project
```
---
## βοΈ Client Configuration
### Claude Desktop
Add to your `claude_desktop_config.json`:
```json
{
"mcpServers": {
"context-tree": {
"command": "uv",
"args": [
"run",
"--directory",
"D:/Repo/mcp-context-tree",
"context-tree"
]
}
}
}
```
### Cursor IDE / Windsurf
Add to `.cursor/mcp.json` or Cursor MCP Settings:
```json
{
"mcpServers": {
"context-tree": {
"command": "uv",
"args": ["run", "--directory", "/absolute/path/to/mcp-context-tree", "context-tree"]
}
}
}
```
### Google Antigravity / Remote SSE Setup
If using network transport (`--transport sse`):
```json
{
"mcpServers": {
"context-tree": {
"url": "http://127.0.0.1:8000/sse"
}
}
}
```
---
## π οΈ MCP Tools Reference
### 1. `index_workspace`
Scans the project directory, computes SHA-256 hashes, applies `.gitignore` rules, and incrementally updates the local ChromaDB vector store.
```json
// Parameters
{
"directory_path": "."
}
// Response
{
"status": "ok",
"workspace": "/path/to/project",
"added": 12,
"modified": 2,
"deleted": 0,
"unchanged": 85,
"indexed_chunks": 340,
"total_in_store": 340
}
```
### 2. `semantic_search`
Executes deep code search across the workspace with live snippet resolution from disk.
```json
// Parameters
{
"query": "how to verify and refresh JWT authentication tokens",
"directory_path": ".",
"limit": 5,
"mode": "hybrid", // "hybrid" | "semantic" | "keyword"
"rerank": true // Optional 2nd-stage Cross-Encoder re-ranking
}
// Response
{
"results": [
{
"file": "src/auth/service.py",
"type": "method",
"class": "AuthService",
"name": "verify_jwt_token",
"start_line": 45,
"end_line": 68,
"score": 0.9624,
"code": "def verify_jwt_token(self, token: str) -> Claims:\n ..."
}
]
}
```
### 3. `find_ast_usages`
Performs true AST-level call-site resolution for functions, methods, and classes, ignoring string literals and comments.
```json
// Parameters
{
"symbol_name": "AuthService.verify_jwt_token",
"directory_path": ".",
"limit": 50
}
// Response
{
"usages": [
{
"file": "src/api/routes.py",
"line": 104,
"preview": "claims = auth_service.verify_jwt_token(token)"
}
]
}
```
### 4. `go_to_definition`
Instantly resolves the exact AST declaration/definition location of a symbol across all 12 supported languages.
```json
// Parameters
{
"symbol_name": "UserRepo.getUser",
"directory_path": ".",
"limit": 20
}
// Response
{
"definitions": [
{
"file": "src/models/User.kt",
"language": "kotlin",
"type": "method",
"name": "getUser",
"class": "UserRepo",
"start_line": 14,
"end_line": 22,
"code": "fun getUser(id: String): User? {\n ...",
"docstring": "/** Retrieve user by identifier */"
}
]
}
```
---
## π¬ Search & Ranking Algorithm
ContextTree MCP uses a **3-Layer Reciprocal Rank Fusion (RRF)** formula to merge dense semantic embeddings, exact lexical matches, and architectural importance:
$$RRF(d) = \frac{w_{vec}}{k + rank_{vec}(d)} + \frac{w_{bm25}}{k + rank_{bm25}(d)} + \frac{w_{graph}}{k + rank_{graph}(d)}$$
Where:
- $k = 60$ (smoothing constant)
- $w_{vec} = 1.0$ (dense semantic similarity via `all-MiniLM-L6-v2`)
- $w_{bm25} = 1.0$ (Robertson-SpΓ€rck Jones BM25 with `camelCase`/`snake_case` tokenization)
- $w_{graph} = 0.5$ (in-degree call frequency boost: heavily referenced core symbols float to the top)
- **Cross-Encoder Layer:** When `rerank=True`, candidate chunks pass through joint self-attention (`cross-encoder/ms-marco-MiniLM-L-6-v2`) for fine-grained semantic scoring.
---
## π Privacy & Security
- **100% Local Execution:** All parsing, embedding generation, and vector indexing happen entirely on your machine.
- **Zero Cloud Network Calls:** Never transmits source code or embeddings to external APIs.
- **Respects Ignore Rules:** Honors root and nested `.gitignore` rules alongside built-in filters for `target/`, `node_modules/`, `bin/`, `obj/`, `.git/`, `.venv/`.
---
## π§ͺ Testing & Quality
ContextTree MCP maintains **100% pass rate** across its test suite and strict linting:
```bash
# Run test suite (60 unit & integration tests)
uv run pytest
# Run linter and formatting check
uv run ruff check .
uv run ruff format --check .
```
---
## π License
Distributed under the **MIT License**. See [LICENSE](LICENSE) for details.
<div align="center">
<sub>Built with β€οΈ for AI engineers and developers. Star β this repository if you find it helpful!</sub>
</div>
TDQS
Scored across 4 tools
Each tool targets a distinct part of code navigation: semantic search vs. exact definitions vs. AST usages vs. index maintenance. There is no realistic confusion between their purposes.
Three of four tools use imperative verb-first names (find_ast_usages, index_workspace, go_to_definition), and all are snake_case. semantic_search breaks the pattern slightly by starting with an adjective instead of a verb, but the naming remains readable and predictable.
With 4 tools, the server is well-scoped and each tool earns its place in the code indexing/search/navigation workflow. There is no redundancy or bloat.
The core lifecycle is covered: index the workspace, search semantically, go to definitions, and find AST usages. Minor gaps such as no explicit index status or reset operation prevent a perfect score, but agents can work around them.