Codebase Copilot
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Codebase CopilotFind where InsufficientStockException is thrown and show the surrounding code"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Codebase Copilot
Make a source repository queryable by AI agents. Codebase Copilot indexes a repository with
language-aware chunking and hybrid (BM25 + vector) retrieval, serves it to any MCP client
(Claude Desktop, Claude Code) as a set of narrow, path-guarded tools, and ships an agentic
code-review CLI that runs a Claude tool-use loop over those same tools and returns
schema-validated findings that point at real file:line locations.
Everything except the two commands that call Claude (ask, review) runs locally with no API
key, and the whole test suite runs offline.
Package | What it does |
Repository walker (honours | |
MCP server on stdio (official | |
| |
Versioned prompt files with front matter; every output records the prompt id, version and hash | |
Golden retrieval cases and an offline runner reporting recall@k, MRR and refusal accuracy | |
A small Spring-style Java order service (plus one TS client) used by tests, evals and demos |
Architecture
flowchart LR
subgraph Ingest
R[Repository] --> W["Walker<br/>.gitignore, size, binary and symlink filters"]
W --> C["Chunker<br/>class / method boundaries<br/>sliding-window fallback"]
end
subgraph Index
C --> B[BM25 index]
C --> E["EmbeddingProvider<br/>local TF-IDF or Voyage AI"]
E --> V["Vector store<br/>cosine, persisted as JSON"]
end
subgraph Retrieve
Q[Query] --> B
Q --> V
B --> F[Reciprocal Rank Fusion]
V --> F
F --> G["Confidence gate<br/>term coverage + similarity"]
end
subgraph Generate
G -->|confident| A["Claude answer<br/>must cite path:line"]
G -->|low confidence| X[Refuse before any API call]
A --> CV[Citation verifier]
end
F --> T["Tool handlers<br/>zod-validated, path-guarded"]
T --> M[MCP server - stdio]
M --> CL[Claude Desktop / Claude Code]
T --> AG["Review agent<br/>Claude tool-use loop"]
AG --> J["submit_review<br/>schema-validated, grounded findings"]The MCP server and the review agent share one tool registry (packages/mcp-server/src/tools.ts):
the server exposes it over the protocol, the agent calls the same handlers in-process.
Related MCP server: Ariadne
Quickstart
Requires Node.js 20 or newer.
npm ci
npm run build
npm test # offline; no API key needed
# Index and search the bundled sample repository (no API key needed)
node packages/agent/dist/cli.js index --root sample-repo
node packages/agent/dist/cli.js search "what happens when stock is insufficient" --root sample-repo --k 4Actual output:
src/main/java/com/example/orders/exception/InsufficientStockException.java:3-3 [class InsufficientStockException] bm25#2 vec#1
src/main/java/com/example/orders/exception/InsufficientStockException.java:5-7 [constructor InsufficientStockException.InsufficientStockException] bm25#1 vec#2
src/test/java/com/example/orders/service/InventoryServiceTest.java:19-23 [method InventoryServiceTest.reserveThrowsWhenStockIsInsufficient] bm25#3 vec#4
src/main/java/com/example/orders/exception/InsufficientStockException.java:1-1 [block] bm25#5 vec#3
confidence: ok; term coverage 67%And for a question the code base cannot answer:
$ node packages/agent/dist/cli.js search "how are payroll taxes withheld" --root sample-repo
confidence: LOW (only 0/3 query terms appear in the retrieved code (missing: payroll, tax, withheld)); term coverage 0%To use Claude (ask and review), copy .env.example or export the key:
export ANTHROPIC_API_KEY=... # never commit it; .env is git-ignored
node packages/agent/dist/cli.js ask "What happens to stock when an order is cancelled?" --root sample-repo
node packages/agent/dist/cli.js review src/main/java/com/example/orders/service/OrderService.java --root sample-repo
node packages/agent/dist/cli.js review --diff main --root . # review your working tree against main
node packages/agent/dist/cli.js review --diff-file change.patch --root .
node packages/agent/dist/cli.js review src/Foo.java --root . --dry-run # print the exact request, no API callVariable | Default | Purpose |
| - | Required for |
|
| Claude model ID |
|
|
|
|
| Server-side refusal fallback ( |
|
|
|
Use it from an MCP client
The server speaks MCP over stdio and indexes the repository at startup (it reuses
<root>/.copilot-index/index.json if present and still fresh; --persist writes it).
Logs go to stderr; stdout carries protocol messages only.
Claude Code - add it from the command line:
claude mcp add codebase-copilot -- node /absolute/path/to/codebase-copilot-mcp/packages/mcp-server/dist/cli.js --root /absolute/path/to/your/repoor commit a project-scoped .mcp.json at the root of the repository you want to query:
{
"mcpServers": {
"codebase-copilot": {
"command": "node",
"args": [
"/absolute/path/to/codebase-copilot-mcp/packages/mcp-server/dist/cli.js",
"--root",
"."
]
}
}
}Claude Desktop - claude_desktop_config.json:
{
"mcpServers": {
"codebase-copilot": {
"command": "node",
"args": [
"/absolute/path/to/codebase-copilot-mcp/packages/mcp-server/dist/cli.js",
"--root",
"/absolute/path/to/your/repo"
]
}
}
}Docker - the image serves whatever is mounted at /workspace (read-only is fine):
docker build -t codebase-copilot-mcp .
docker run -i --rm -v "$PWD:/workspace:ro" codebase-copilot-mcp{
"mcpServers": {
"codebase-copilot": {
"command": "docker",
"args": ["run", "-i", "--rm", "-v", "/absolute/path/to/your/repo:/workspace:ro", "codebase-copilot-mcp"]
}
}
}Tools
Tool | Input | Returns |
|
| Ranked chunks with |
|
| Numbered lines (max 400 per call) from an indexed file |
|
| Signature, doc comment, code, enclosing class, members (for classes), reference count; suggestions when not found |
|
| Whole-word occurrences as |
|
| A JUnit 5 + Mockito or Jest skeleton derived from the real signature, collaborators and thrown exceptions |
resource | - | Every indexed file with language, size, line count and chunk count |
suggest_tests for OrderService.cancelOrder produces (abridged):
@ExtendWith(MockitoExtension.class)
class OrderServiceCancelOrderTest {
@Mock
private OrderRepository orderRepository;
// ... @Mock InventoryService, PaymentService, DiscountPolicy
@InjectMocks
private OrderService orderService;
@Test
void cancelOrder_returnsExpectedResult() {
Long orderId = 1L;
// Arrange: stub the collaborators this path uses: orderRepository.findById, orderRepository.save, inventoryService.release, paymentService.refundOrder
Order result = orderService.cancelOrder(orderId);
assertNotNull(result);
}
@Test
void cancelOrder_throwsOrderNotFoundException() {
Long orderId = 1L;
// Arrange: set up the condition that makes cancelOrder throw OrderNotFoundException
assertThrows(OrderNotFoundException.class, () -> orderService.cancelOrder(orderId));
}
}The skeleton compiles against the real signature and imports; the assertions are starting points for a human to finish, not finished tests.
AI-assisted code review
copilot review builds a review target (numbered file contents, or a unified diff annotated with
new-file line numbers), then runs a manual Claude tool-use loop:
The model reads the target and calls the codebase tools to gather context (callers, collaborators, related validation).
All tool calls from one turn run in parallel and go back in a single message; tool errors (bad input, blocked paths) are returned as
is_errorresults so the model can correct itself.The run ends when the model calls
submit_reviewwith input that passes the zod schema below. Invalid submissions are sent back with the validation errors.Each finding is grounded: the file must be under review (or indexed) and the line must exist. Findings that fail are reported separately and never presented as results.
{
summary: string,
findings: Array<{
severity: "critical" | "high" | "medium" | "low" | "info",
file: string, // repository-relative path
line: number, // 1-based, verified to exist
issue: string, // what is wrong and its concrete consequence
suggestion: string, // the change that fixes it
category?: "correctness" | "security" | "data-integrity" | "error-handling"
| "concurrency" | "performance" | "testing" | "maintainability"
}>
}Example of the text output format. This was rendered by the real ReviewAgent and
formatOutcome code, but the model turn was scripted (a fake client, as in
agent.test.ts), so token counts are zero and the findings were written by hand to show the
format, including a finding the grounding check discards:
Review completed - 2 model calls, 2 tool calls, 0 in / 0 out tokens (cache read 0, cache write 0), est. $0.0000
Scripted example: cancelOrder has no state guard, so repeated or late cancellations refund and restock again.
[HIGH] src/main/java/com/example/orders/service/OrderService.java:69 (correctness)
Issue: cancelOrder never checks the current status, so cancelling a SHIPPED or already CANCELLED order refunds the payment again and returns stock that has left the warehouse.
Suggestion: Allow cancellation only from PENDING or PAID and throw an IllegalStateException otherwise.
1 finding(s) discarded because their location could not be verified:
- src/main/java/com/example/orders/service/OrderService.java:999: src/main/java/com/example/orders/service/OrderService.java has 89 lines; line 999 does not exist--format json prints the full outcome (status, report, rejected findings, per-tool call log,
token usage, estimated cost and the prompt versions used). Structured JSON event logs
(llm_call, llm_retry, tool_call, review_finished) go to stderr.
The sample repository contains a few deliberate problems for the reviewer to find: string-built
SQL in JdbcOrderSearchRepository, double arithmetic for money in DiscountPolicy, and a
cancelOrder with no status check.
Evaluation
npm run build && npm run eval runs the golden set in evals/golden.json
against the sample repository with the local TF-IDF embeddings: no network, fully deterministic.
Relevance is judged at file level, or at file#symbol level for cases about one method. Results
from the current code:
Mode | Recall@1 | Recall@3 | Recall@5 | MRR |
bm25 | 63.6% | 81.8% | 86.4% | 0.748 |
vector | 63.6% | 95.5% | 100.0% | 0.758 |
hybrid | 63.6% | 86.4% | 100.0% | 0.745 |
Confidence gate (hybrid): 3/3 unanswerable questions refused, 0/11 answerable questions wrongly refused.
How to read this honestly:
The set is small (11 answerable, 3 unanswerable questions over a 22-file corpus), so one case moves recall by about 9 points. It is a regression check, not a benchmark. CI fails if hybrid recall@5 drops below 90%.
On this corpus the vector leg alone is slightly better than the fused result at recall@3. The local "vector" leg is TF-IDF with character trigrams, which is still lexical; fusion with BM25 helps most when the two legs fail differently, which is the case a neural embedding model (
COPILOT_EMBEDDINGS=voyage) is meant to create. That configuration has not been evaluated here because it needs an API key.The unanswerable questions (payroll taxes, Kafka topics, JWT key rotation) share no vocabulary with the corpus, which makes them easy for a term-coverage gate. Questions that reuse the corpus vocabulary but ask for something absent are harder and are not in this set.
Design decisions
Hybrid retrieval with Reciprocal Rank Fusion. Code search mixes two kinds of queries: exact
identifiers (JdbcOrderSearchRepository, OrderNotFoundException) where BM25 is hard to beat,
and descriptive questions ("where is stock released when an order is cancelled") where
embeddings help. RRF merges the two rankings using ranks only, so BM25 scores and cosine
similarities never need to be normalized onto one scale. Each leg nominates max(30, 3k)
candidates before fusion.
Chunking on code structure. A small lexer masks strings, char literals, text blocks, template
literals and comments, so brace matching is not fooled by "{" or // }. Methods, constructors
and functions become their own chunks together with their Javadoc/JSDoc and annotations, and
carry symbol, container and signature metadata; a class chunk holds the declaration and
fields; imports and stray code become block chunks, so every non-trivial line is covered. Chunks
longer than 80 lines, and files in other languages, fall back to overlapping 40-line windows.
The indexed text for each chunk is prefixed with its path and symbol names, which is often the
strongest signal. Tokenization splits camelCase/snake_case while also keeping the whole
identifier.
Grounding and refusal. Refusal happens in layers:
Before any tokens are spent: retrieval reports term coverage (share of query terms present in the results) and the best cosine similarity; below the thresholds,
askrefuses without calling Claude.In the prompt: the model gets numbered sources and a single machine-checkable way out,
INSUFFICIENT_CONTEXT: ..., which is surfaced as a refusal.After generation: every
[path:start-end]citation is checked against the retrieved sources' line ranges. An answer with no verifiable citation is markedungrounded; invalid citations are listed, never silently kept.For reviews: findings must name an existing file and line or they are discarded.
A stop_reason of refusal is handled explicitly, and the server-side refusal fallback
(fallbacks: "default") is enabled by default and can be turned off.
Cost control.
Refuse early: low-confidence questions never reach the API.
Prompt caching: the review system prompt carries an explicit cache breakpoint (tools are rendered before the system prompt, so they are cached with it), and top-level automatic caching covers the growing conversation. The history is append-only (assistant turns are sent back exactly as received), so earlier turns stay cacheable. System prompts contain no per-request data for the same reason.
Bounded loops: a max-iteration guard (default 8 model calls), a single nudge if the model answers in prose instead of submitting, and hard limits on tool output size (400 lines per snippet, 60 lines per search hit, 200 references).
Effort is configurable per command (
highfor review,mediumfor Q&A).Every call logs input, output, cache-write and cache-read tokens with latency; cost is estimated per completed task with the published per-million-token prices (cache writes at 1.25x, reads at 0.1x the input price). The estimate uses the requested model's price, so it is approximate if a fallback model served a turn.
Retries are owned in one place: the SDK's built-in retries are disabled and
withRetryretries 429, 5xx (including 529) and connection errors with exponential backoff and jitter, honoursretry-after, and never retries other 4xx errors.
Security. Model-supplied paths are untrusted input. resolveIndexedFile rejects absolute
paths, drive letters, NUL bytes and .. escapes, re-checks containment after realpath (so a
symlink cannot point outside the root), only serves files that are in the index (so .env,
.git/ and ignored files are unreachable even inside the root) and enforces a size cap. The
walker never follows symlinks and always ignores .env*, .git/ and node_modules/. Tool inputs
are validated with zod before any handler runs, git diff is invoked with execFile and a
validated revision (no shell), and prompts tell the model that code and comments are data, not
instructions.
Prompt engineering. Prompts live in prompts/ as versioned files
(review.system.v1.md, review.task.v1.md, answer.system.v1.md) with front matter. The loader
takes the highest version unless one is pinned, and the prompt id, version and a content hash are
recorded with every answer and review, so an output can be traced to the exact instructions that
produced it. See prompts/README.md for the conventions.
Testing and CI
136 Jest tests (ts-jest) across the three packages, all offline: chunker and lexer edge cases, BM25, RRF, the vector store, TF-IDF and Voyage embeddings (mocked
fetch), retry and backoff, the path guard (including symlink escapes, skipped on Windows without symlink privileges), every MCP tool handler called directly, the MCP server over an in-memory transport, and the review agent driven by a scripted fake Anthropic client (tool loop, parallel results, invalid submissions, refusal,pause_turn,max_tokens, 429/529 retries, the iteration guard, grounding).Coverage thresholds are enforced in
jest.config.js(current: about 95% of lines).GitHub Actions runs lint (typecheck, ESLint, Prettier), tests with coverage, the build, the eval with a recall gate, a stdio smoke test that spawns the built server and calls it as an MCP client, CLI smoke tests (
search,review --dry-run), and builds the Docker image and runs the same stdio smoke test against the container. Node 20 and 22.
Limitations
The chunker is a lexer plus heuristics, not a parser. It handles common Java and TypeScript well (see
chunker.test.ts) but unusual formatting can produce a coarser chunk; oversized chunks still fall back to windows, so nothing is lost.find_referencesis lexical: it finds every whole-word occurrence, including same-named methods on unrelated classes and mentions in strings.The vector store is brute-force and in memory: fine for a repository of tens of thousands of chunks, not for a monorepo of millions. The
VectorStoreinterface is where a real vector database would plug in.The live Claude paths (
ask,review) are covered by tests with a scripted client; their output quality depends on the model and has not been measured here.
Project layout
packages/
rag/ src/{ingest,chunking,retrieval,store,embeddings,llm,answer,eval}
mcp-server/ src/{tools,security,symbols,testgen,server,cli}.ts
agent/ src/review/{agent,targets,schema,format}.ts, src/cli.ts
prompts/ versioned prompt files
evals/ golden.json, run-evals.mjs
scripts/ mcp-smoke.mjs
sample-repo/ Java/TS fixture
docs/ ai-assisted-workflow.mdSee CLAUDE.md for contributor rules and docs/ai-assisted-workflow.md
for how AI tools were used to build this repository and how their output was checked.
License
MIT (c) rkunameneni
This server cannot be deployed
Maintenance
Related MCP Connectors
Ask a codebase what calls what: search, blast radius, paths between symbols, and diffs.
Search GitHub, npm, PyPI, StackOverflow, ArXiv from one MCP — built for coding agents.
The Cortex MCP server provides read-only access to real-time engineering context from the Cortex developer portal, allowing AI coding assistants to answer natural language questions about your organization's catalog (microservices, libraries, domains, teams, infrastructure), scorecards (engineering standards and best practices), initiatives (goals and deadlines), and Engineering Intelligence metrics. It includes tools for querying documentation, tracking personal entities, and accessing AI-assisted insights across the entire Cortex ecosystem.
Code intelligence for coding agents: semantic, AST, graph, and full-text search. 279+ languages.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceAgent-safe code retrieval MCP server that indexes repositories and provides semantic search, file navigation, call graph analysis, and bounded file reading tools for coding agents.4,370,901 npm3AGPL 3.0
- AlicenseNot gradedqualityCmaintenanceSemantic code index and gatekeeper that exposes 14 read-only MCP tools for AI agents, enabling symbol search, definition lookup, reference finding, and impact analysis via static analysis of codebases.35 npmMIT
- AlicenseNot gradedqualityAmaintenanceProvides coding agents with searchable codebase context through an MCP server, enabling hybrid BM25 and semantic search, symbol graph navigation, and dependency mapping over an incrementally maintained repository index.247 PyPI84Apache 2.0
- FlicenseNot gradedqualityCmaintenanceExposes code search and file reading tools over the Model Context Protocol, enabling any MCP-compatible client to query a codebase with natural language.-