Codebase Intelligence MCP Server
by S-naruka
README.md
#WIP (Multiple known issues)
# Codebase Intelligence MCP Server
An stdio MCP server that indexes a Python codebase into a local SQLite graph. It offers hybrid code search, fast file context, lazy explanations, dependency traces, and version-aware symbol history.
It is a local, single-user tool—not a language server, hosted service, background crawler, or full type-inference engine.
## Status
Phases 1–9 are implemented: schema and parser, graph traversal, search, lazy LLM summaries, versioning, seven MCP tools, CLI indexing, tests, and documentation.
## Prerequisites
- Python 3.11+
- [uv](https://docs.astral.sh/uv/)
- Git (for commit-aware version tracking)
- [Ollama](https://ollama.com/download) for indexing/search and local summaries
- Node.js for MCP Inspector
Pull the configured local models:
```bash
ollama pull qwen3-embedding:0.6b
ollama pull qwen2.5-coder:7b
```
## Setup and first index
```bash
cd "d:\Projects\MCP Project\codebase-intelligence-mcp"
uv sync
copy .env.example .env
# Set REPO_PATH in .env, if desired.
uv run codebase-intel index "d:\path\to\python-repo"
```
To change embedding models safely:
```bash
uv run codebase-intel reindex-embeddings "d:\path\to\python-repo"
```
Start the MCP server after setting `REPO_PATH` and ensuring Ollama is running:
```bash
uv run python src/server.py
```
## MCP tools
| Tool | Example | Purpose |
|---|---|---|
| `search_codebase` | `{"query":"parse Python files", "top_k":5}` | Hybrid vector + BM25 retrieval with centrality reranking. |
| `get_file_context` | `{"file_path":"src/parser/ast_parser.py"}` | Instant template summary, imports, and signatures. |
| `explain_function` | `{"symbol_name":"parse_file"}` | Cached behavioural summary plus live callers/callees. |
| `trace_dependencies` | `{"symbol_name":"parse_file", "depth":2}` | Breadth-first caller/callee traversal. |
| `detect_changes` | `{"since_version":"last"}` | Git-backed changed-file parsing and version diffs. |
| `get_version_history` | `{"symbol_name":"parse_file"}` | Added/removed/renamed/modified history, including rename lineage. |
| `get_codebase_health` | `{}` | Index counts, lazy-summary coverage, tombstones, and renames. |
## Configuration
See [.env.example](.env.example). Key settings are `OLLAMA_HOST`, `SUMMARY_MODEL`, `EMBEDDING_MODEL`, `USE_REMOTE_SUMMARIES`, `GEMINI_API_KEY`, `DATABASE_PATH`, and `REPO_PATH`.
## Known limitations
- Call-graph heuristic: dynamic dispatch is not resolved. Only unambiguous same-file names become call edges.
- Rename heuristic: renames require an exact matching content hash; an edit plus rename appears as removed and added.
- Concurrency: SQLite WAL plus a 5-second busy timeout is suitable for local use, not heavy simultaneous multi-process writes.
- sqlite-vec scaling: brute-force vector search is comfortable to roughly 50K–100K symbols; it has no ANN index.
## Testing
```bash
uv run pytest tests/ -v
npx @modelcontextprotocol/inspector uv run python src/server.py
```
The repository also contains phase exit scripts under `scripts/`. The CLI and server require local Ollama models; the automated unit tests mock provider calls.
## Future work
- Background crawler and cross-process coordination
- LSP-assisted call resolution and refactor-aware renames
- PageRank or impact-scoped centrality recomputation
- Token-budget-aware response packing and staleness detection
- More languages and ANN vector indexing
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues