repomind
Allows analyzing and querying public GitHub repositories, providing tools to search code, read files, list symbols, and get repository statistics.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@repomindHow is authentication implemented in honojs/hono?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
RepoMind — Codebase Intelligence Agent
▶ Live demo: https://repomind-teal.vercel.app · MCP endpoint · CI: typecheck + tests + eval gate + build
The live demo boots with a pre-indexed sample service (
demo/acme-service) so you can ask questions immediately — try "How is authentication implemented?" Paste any public repo (e.g.honojs/hono) to index your own.
Point it at any public GitHub repo. It ingests the code with AST-aware chunking, indexes it with hybrid retrieval (dense vectors + lexical BM25 fused with Reciprocal Rank Fusion, in one SQL query over pgvector), and lets you chat with an agentic tool loop that answers architecture, "where is X", and how-does-this-work questions — every claim carrying a clickable citation back to the exact file and line.
The same repo tools are exposed as an MCP server over Streamable HTTP, so Claude Desktop / Cursor can use the deployed backend directly. Retrieval quality is defended by an eval harness (recall@k · MRR · nDCG) wired as a CI regression gate.
Runs with zero API keys. Missing keys degrade gracefully — an in-process Postgres (PGlite + pgvector) replaces Neon, and a deterministic local embedding + extractive answerer replace OpenAI. Add keys to upgrade quality, not to boot. This is why the CI suite and the Vercel demo work with no secrets.
Why this project
It's built to show the surface area of a senior AI/ML engineer, end to end:
Capability | What's actually implemented |
Retrieval engineering | Dense (pgvector HNSW) + lexical (Postgres |
RAG quality technique | Contextual Retrieval (Anthropic's method): each chunk is enriched with an LLM-generated situating summary before embedding, which lifts recall. Generated once at ingest. |
AST-aware chunking | Splits code on real symbol boundaries (functions/classes/types) across TS/JS/Python/Go/Rust/Java/…, so chunks are semantically whole and citations name real symbols. |
Agentic tool use | A bounded agent loop (AI SDK v5) with |
MCP | The exact same tools re-exposed as a Model Context Protocol server — one implementation, two transports. |
Evaluation | Golden Q&A set scored with recall@k / MRR / nDCG, an ablation (dense-only vs lexical-only vs hybrid), and a CI gate that fails the build on regression. |
Production concerns | Prompt-injection screening, semantic cache (embedding-similarity answer reuse), rate limiting, and token/cost tracing on a live observability dashboard. |
Data engineering | GitHub tarball streaming ingest (one request, not one-per-file), content-hash incremental reindex (only changed chunks re-embed; deleted chunks pruned). |
Dual backend | Identical SQL over Neon serverless Postgres (prod) and PGlite (local/CI). Backend-agnostic app code. |
Related MCP server: Code Understanding MCP Server
Architecture
flowchart LR
subgraph Ingest
GH[GitHub tarball] --> CH[AST chunker<br/>symbol boundaries]
CH --> CTX[Contextual<br/>enrichment]
CTX --> EMB[Embeddings<br/>OpenAI / local]
EMB --> PG[(pgvector + tsvector<br/>Neon / PGlite)]
end
subgraph Query
Q[Question] --> GUARD[Injection screen<br/>+ rate limit]
GUARD --> CACHE{Semantic<br/>cache?}
CACHE -- hit --> ANS
CACHE -- miss --> HYB[Hybrid retrieve<br/>dense + lexical → RRF]
HYB --> RR[Rerank<br/>LLM / fusion order]
RR --> AGENT[Agent loop<br/>tools + streaming]
AGENT --> ANS[Cited answer]
ANS --> LOG[(Telemetry)]
end
PG --- HYB
AGENT -. same tools .-> MCP[[MCP server<br/>/api/mcp]]Retrieval benchmark
From npm run eval (hermetic: fixture repo, local hashing embeddings, PGlite). The ablation is the point — hybrid RRF fusion beats either arm alone on every metric:
Configuration | Recall@5 | MRR | nDCG@10 | Hit@5 |
dense-only | 96.4% | 0.929 | 91.8% | 100% |
lexical-only | 96.4% | 0.893 | 88.5% | 100% |
hybrid (RRF) | 96.4% | 1.000 | 95.5% | 100% |
14 golden questions spanning paraphrase (dense) and exact-token (lexical) queries. The CI gate fails the build if hybrid drops below recall@5 0.75 / MRR 0.60 / nDCG 0.65. With real OpenAI embeddings the absolute numbers rise further; these are a regression floor, not a ceiling.
Run it locally
npm install
npm run dev # http://localhost:3000 — works with no keysThen paste a repo like tiangolo/fastapi (or click a sample) and ask questions.
npm test # 22 tests: chunker, embeddings, metrics, + PGlite integration
npm run eval # print the retrieval benchmark, write evals/results.json
npm run eval -- --ci # same, but exit non-zero on regression (used in CI)
npm run typecheck # strict TS, no errors
npm run build # production buildOptional configuration (.env.local)
Everything is optional — see .env.example.
Var | Effect |
| Switches embeddings to |
| Neon serverless Postgres (needs the |
| Raises GitHub rate limits and allows private repos. |
| If set, the MCP endpoint requires |
Use it from Claude Desktop / Cursor (MCP)
The deployment is a live MCP server. Add to your client config:
{
"mcpServers": {
"repomind": { "url": "https://<your-deployment>.vercel.app/api/mcp" }
}
}Tools exposed: list_repos, search_code, read_file, list_symbols, repo_stats.
Deploy to Vercel
Push to GitHub, import the repo in Vercel (framework auto-detected).
It deploys and runs with no env vars (PGlite on
/tmp, local models).For durable, multi-instance storage and real LLM answers, add
DATABASE_URL(Neon) andOPENAI_API_KEYin Project → Settings → Environment Variables, then redeploy.
Layout
src/lib/db/ dual-backend Postgres (client, schema, vector helpers)
src/lib/ingest/ github tarball stream · AST chunker · contextual enrichment · pipeline
src/lib/retrieval/ hybrid RRF search (one SQL query) · reranker
src/lib/agent/ tools · agent engine · guardrails · semantic cache · rate limit
src/lib/eval/ IR metrics · harness (with ablation)
src/lib/obs/ token/cost accounting · telemetry
src/app/api/ chat · ingest · mcp · repos · stats · observability · eval · file
src/components/ premium streaming UI (chat, citations, code drawer, dashboards)
evals/ fixture repo · golden set · committed results.json
test/ vitest suites (unit + PGlite integration)Built with Next.js 16, AI SDK v5, @modelcontextprotocol/server + mcp-handler, @neondatabase/serverless, @electric-sql/pglite + pgvector.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- AlicenseBqualityDmaintenanceAn MCP server for gitingest. It allows MCP clients like Claude Desktop, Cursor, Cline etc to quickly extract information about Github repositories including repository summaries, project directory structure, file contents, etc3136MIT
- AlicenseCqualityDmaintenanceAn MCP server that analyzes local or remote GitHub repositories, providing intelligent code context and structure to AI coding assistants.1013MIT
- AlicenseDqualityDmaintenanceA lightweight MCP server for bringing GitHub repositories into context for large language models, enabling repository analysis, file access, and search without local cloning.496Apache 2.0
- AlicenseBqualityCmaintenanceA minimal MCP server for GitHub that enables browsing repos, reading code, searching, and getting insights through natural language via Claude, Cursor, or any MCP client.7MIT
Related MCP Connectors
An MCP server that gives your AI access to the source code and docs of all public github repos
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
Repo intel for AI coding agents: overview, PRs, contributors, hot files, CI, deps. Remote MCP.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/shashvat-singham/repomind'
If you have feedback or need assistance with the MCP directory API, please join our Discord server