Codebase Intelligence MCP Server
#WIP (Multiple known issues)
Codebase Intelligence MCP Server
An stdio MCP server that indexes a Python codebase into a local SQLite graph. It offers hybrid code search, fast file context, lazy explanations, dependency traces, and version-aware symbol history.
It is a local, single-user tool—not a language server, hosted service, background crawler, or full type-inference engine.
Status
Phases 1–9 are implemented: schema and parser, graph traversal, search, lazy LLM summaries, versioning, seven MCP tools, CLI indexing, tests, and documentation.
Prerequisites
Python 3.11+
Git (for commit-aware version tracking)
Ollama for indexing/search and local summaries
Node.js for MCP Inspector
Pull the configured local models:
ollama pull qwen3-embedding:0.6b
ollama pull qwen2.5-coder:7bSetup and first index
cd "d:\Projects\MCP Project\codebase-intelligence-mcp"
uv sync
copy .env.example .env
# Set REPO_PATH in .env, if desired.
uv run codebase-intel index "d:\path\to\python-repo"To change embedding models safely:
uv run codebase-intel reindex-embeddings "d:\path\to\python-repo"Start the MCP server after setting REPO_PATH and ensuring Ollama is running:
uv run python src/server.pyMCP tools
Tool | Example | Purpose |
|
| Hybrid vector + BM25 retrieval with centrality reranking. |
|
| Instant template summary, imports, and signatures. |
|
| Cached behavioural summary plus live callers/callees. |
|
| Breadth-first caller/callee traversal. |
|
| Git-backed changed-file parsing and version diffs. |
|
| Added/removed/renamed/modified history, including rename lineage. |
|
| Index counts, lazy-summary coverage, tombstones, and renames. |
Configuration
See .env.example. Key settings are OLLAMA_HOST, SUMMARY_MODEL, EMBEDDING_MODEL, USE_REMOTE_SUMMARIES, GEMINI_API_KEY, DATABASE_PATH, and REPO_PATH.
Known limitations
Call-graph heuristic: dynamic dispatch is not resolved. Only unambiguous same-file names become call edges.
Rename heuristic: renames require an exact matching content hash; an edit plus rename appears as removed and added.
Concurrency: SQLite WAL plus a 5-second busy timeout is suitable for local use, not heavy simultaneous multi-process writes.
sqlite-vec scaling: brute-force vector search is comfortable to roughly 50K–100K symbols; it has no ANN index.
Testing
uv run pytest tests/ -v
npx @modelcontextprotocol/inspector uv run python src/server.pyThe repository also contains phase exit scripts under scripts/. The CLI and server require local Ollama models; the automated unit tests mock provider calls.
Future work
Background crawler and cross-process coordination
LSP-assisted call resolution and refactor-aware renames
PageRank or impact-scoped centrality recomputation
Token-budget-aware response packing and staleness detection
More languages and ANN vector indexing
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/S-naruka/MCP'
If you have feedback or need assistance with the MCP directory API, please join our Discord server