Repo-Aware Code RAG MCP Server
by ManavK003
README.md
# Repo-Aware Code RAG Assistant
[](https://github.com/ManavK003/repo-aware-code-rag/actions/workflows/ci.yml)
> **Note:** This repository is the **public demonstration version** of a private, live system under continuous development. It implements the same architecture end-to-end — AST-aware chunking, hybrid retrieval, the FastAPI service, the MCP server, and the evaluation harness — in a fully local, zero-cloud-account mode. The production deployment (Azure AI Search, Databricks ingestion jobs, MLflow tracking) is private; its integration seams are visible here as thin, documented adapters.
A retrieval-augmented assistant for codebases that answers "where is X and how does it work?" with **exact `file:line` citations** — built on the observation that code retrieval fails differently than prose retrieval, and needs code-specific treatment at every stage.
```
┌──────────────────────────────────────────────┐
repo on disk ──────▶│ AST-aware chunker (chunking.py) │
│ functions/classes with exact line spans; │
│ windows only as fallback │
└──────────────┬───────────────────────────────┘
▼
┌──────────────────────────────────────────────┐
│ Index (index.py) │
│ BM25 w/ identifier field-boost + dense │
│ local: numpy+rank_bm25 · prod: Azure AI Search │
└──────────────┬───────────────────────────────┘
▼
┌──────────────────────────────────────────────┐
│ Retrieval (retrieval.py) │
│ query expansion → RRF fusion → rank priors │
│ → optional cross-encoder rerank │
└──────┬───────────────────────┬───────────────┘
▼ ▼
FastAPI service (api.py) MCP server (mcp_server.py)
/ingest /search /ask search_code · read_span
— retrieval as a live tool
```
## Quickstart (fully local, no keys)
```bash
pip install -e .
python -m repo_rag.ingest . # index this repo on itself
repo-rag-api # http://localhost:8000/docs
```
Real output of the retriever, dogfooding on this repository:
```
$ search: 'where is reciprocal rank fusion implemented'
src/repo_rag/retrieval.py:64-73 (function reciprocal_rank_fusion, via bm25+dense)
src/repo_rag/eval.py:45-68 (function evaluate, via bm25+dense)
src/repo_rag/retrieval.py:109-125 (function Retriever._prior, via bm25+dense)
$ search: 'which MCP tool reads exact source lines for a citation'
src/repo_rag/mcp_server.py:72-82 (function read_span, via bm25+dense)
src/repo_rag/mcp_server.py:1-16 (module mcp_server docstring, via bm25+dense)
src/repo_rag/chunking.py:83-98 (function chunk_repo, via bm25+dense)
```
Ask over HTTP:
```bash
curl -s localhost:8000/search -X POST -H 'content-type: application/json' \
-d '{"query": "how are python files split into chunks", "k": 3}' | jq .
```
`/ask` returns a synthesized answer: extractive with citations by default, or
LLM-written (still citation-constrained) when `ANTHROPIC_API_KEY` is set.
## MCP: retrieval as a live tool, not a hardcoded pipeline
Instead of one fixed ask→retrieve→generate chain, the index is exposed over
the **Model Context Protocol**, so any MCP client — Claude Desktop, Claude
Code, IDE agents — decides *when* and *how* to query it, iteratively,
mid-conversation.
```bash
pip install ".[mcp]"
```
Claude Desktop config (`claude_desktop_config.json`):
```json
{
"mcpServers": {
"repo-rag": {
"command": "repo-rag-mcp",
"env": { "REPO_RAG_INDEX_DIR": "/absolute/path/to/.ragindex" }
}
}
}
```
Tools exposed: `search_code(query, k)`, `read_span(path, start, end)`, `index_stats()`.
## Evaluation
The harness measures **recall@k** and **MRR** against gold `file:line` spans
(hit = line-overlap in the same file), and logs to MLflow when
`MLFLOW_TRACKING_URI` is set. A RAGAS faithfulness pass over generated
answers is available behind `--ragas` (needs the `[eval]` extra and an LLM key).
```bash
repo-rag-eval eval/sample_eval.jsonl --k 5
# recall_at_5: 0.8333
# mrr: 0.6167
# n_questions: 6
```
(The one miss is instructive: this README quotes the demo queries, so once
indexed it outcompetes the implementation for one question - the same
docs-vs-code tension the rank priors exist to manage.)
Those numbers are the **illustrative sample set in this repo** (6
self-referential questions, offline hash embedder, no rerank). The headline
metrics for this project — **recall@5 lifted 61% → 84% (0.79 MRR) with RAGAS
faithfulness 0.82** — were measured on the full evaluation set of the private
live deployment over a **~2K-file codebase**, using the same harness with
production embeddings and reranking enabled. The methodology here is the
methodology there; run it on your own repo with your own eval set.
Ablations are one flag away: `--no-expand` disables query expansion;
`REPO_RAG_RERANK=true` enables the cross-encoder (with the `[ml]` extra).
## Design notes (the interesting 20%)
- **AST chunking over fixed windows** (`chunking.py`) — windows slice through
function bodies, poisoning embeddings and producing citations that start
mid-`def`. AST boundaries keep units intact, give every chunk a *name*, and
make `path:start-end` land exactly where a developer would open the file.
Oversized definitions window *internally*, carrying the parent's name.
- **RRF over weighted score sums** (`retrieval.py`) — BM25 scores and cosine
similarities live on incomparable scales; rank fusion is scale-free and
needs no corpus-specific tuning.
- **Identifier field-boosting in BM25** (`index.py`) — in code search, a chunk
whose *name* matches the query nearly always beats prose that merely
mentions the concept. Implemented as token repetition inside BM25 -
together with the rank priors, the single biggest sample-eval lift.
- **Weak rank priors** — implementations over tests, definitions over module
prose. Deliberately weak multipliers, so strong matches in tests still surface.
- **Hermetic by default** — the deterministic hash embedder keeps tests, CI,
and the demo fully offline; `sbert`/`azure` backends are a config switch.
## Production architecture (private deployment)
- **Azure AI Search** serves hybrid BM25 + vector ranking over the same chunk
schema (`index.AzureAISearchIndex` is the adapter seam).
- **Databricks** runs `repo_rag.ingest` as a scheduled job against repository
checkouts, feeding the indexer.
- **MLflow** tracks every eval run across chunking/retrieval configurations.
- **Docker → Azure**: `docker build -t repo-rag . && az containerapp up --name repo-rag --source .`
## Roadmap
- [x] AST-aware chunking with exact line spans
- [x] Hybrid BM25 + dense retrieval with RRF
- [x] Query expansion; optional cross-encoder rerank
- [x] FastAPI service with citation-formatted responses
- [x] MCP server (stdio): `search_code`, `read_span`, `index_stats`
- [x] Eval harness: recall@k, MRR, MLflow logging, RAGAS hook
- [ ] Incremental re-indexing on file change (watch mode)
- [ ] Tree-sitter chunking for TypeScript/Go/Java (beyond line windows)
- [ ] MCP resources: expose indexed files as browsable resources
- [ ] Eval dashboard comparing configs across runs
## License
MIT © 2026 Manav Kanaganapalli