repo-semantic-search
by shlomihassan
README.md
# repo-semantic-search
Semantic code search for any local repo, available to Claude Code as an MCP
tool instead of grep.
## What this is
A custom pipeline, built from scratch: CocoIndex chunks and embeds a repo's
files (tree-sitter-aware chunking, Ollama for local embeddings), the vectors
land in Postgres/pgvector, and a small MCP server (`repo_index.mcp_server`)
exposes semantic search over them as Claude Code tools. A git `post-commit`
hook keeps each registered repo's index in sync automatically.
An earlier version of this README described adopting a third-party tool,
[cocoindex-code](https://github.com/cocoindex-io/cocoindex-code), instead of
building this. That path was abandoned in favor of the custom
Postgres/pgvector pipeline described below, which is now built, registered
with Claude Code, and verified end-to-end against a real repo.
## Architecture
```mermaid
flowchart TD
subgraph Indexing["Indexing (write path)"]
Repo["Any registered git repo"] -->|git commit| Hook["post-commit hook<br/>nohup, non-blocking"]
Hook --> CLI["repo-index CLI<br/>add / sync / status / install-hook / init"]
CLI --> Registry["registry.py<br/>repos table"]
CLI --> Flow["flow.py<br/>CocoIndex pipeline"]
Flow -->|chunk + embed| Ollama["Ollama<br/>nomic-embed-text"]
Flow -->|upsert rows, repoindex role| PG[("Postgres + pgvector<br/>code_chunks table")]
Registry -->|repoindex role| PG
end
subgraph Querying["Querying (read path)"]
Claude["Claude Code"] -->|semantic_search<br/>list_indexed_repos| MCP["mcp_server.py<br/>MCP server"]
MCP -->|embed query| Ollama
MCP -->|SELECT only, repoindex_ro role| PG
end
```
Two independent paths sharing one Postgres database: indexing (triggered by
commits, writes via the read-write `repoindex` role) and querying (triggered
by Claude Code, reads via the read-only `repoindex_ro` role — the MCP server
has no write path at all).
## Setup
Prerequisites, in order — `repo-index init` (below) will fail with a raw
connection-refused traceback if Postgres isn't running yet.
1. **Create the venv and install dependencies:**
```bash
python3 -m venv .venv
.venv/bin/pip install --group dev -e .
```
Note: `pip install -e '.[dev]'` silently does **not** install the dev
dependencies for this project's `pyproject.toml` — dev deps live in a
PEP 735 `[dependency-groups]` table, not an extra. Always use
`--group dev` as shown above.
2. **Start Postgres (with pgvector):**
```bash
docker compose -f docker/postgres-compose.yml up -d
```
3. **Install Ollama and pull the embedding model:**
```bash
brew install ollama
brew services start ollama
ollama pull nomic-embed-text
```
4. **(Optional) Customize config:** copy `.env.example` to `.env` and edit
as needed. Defaults assume the local Postgres/Ollama setup above.
`TEST_DATABASE_URL` (defaults to `repoindex_test` on the same Postgres
instance) is used only by the test suite (`tests/conftest.py`), which
truncates its tables between runs — keep it pointed at a separate
database from `DATABASE_URL` so tests never touch real registered-repo
data.
### Components
- `repo_index/settings.py` — loads Postgres/Ollama config from env vars
(`DATABASE_URL`, `READONLY_DATABASE_URL`, `OLLAMA_API_BASE`,
`OLLAMA_EMBED_MODEL`), with sane localhost defaults.
- `repo_index/registry.py` — the `repos` table: which repos are registered,
their filesystem path, and last-synced commit/timestamp.
- `repo_index/flow.py` — the CocoIndex flow that chunks files, embeds them
via Ollama, and writes rows into the shared `code_chunks` pgvector table.
- `repo_index/sync.py` — orchestrates a sync run for one registered repo
(resolve HEAD commit, run the flow, update the registry).
- `repo_index/hooks.py` + `install-hook` CLI command — installs a
`post-commit` git hook that re-syncs a repo's index in the background
after every commit, without blocking or failing the commit itself.
- `repo_index/cli.py` — the `repo-index` command-line tool (`add`, `sync`,
`status`, `install-hook`, `init`).
- `repo_index/mcp_server.py` — the MCP server, exposing `semantic_search`
and `list_indexed_repos` tools.
## Adding a new repo to the index
```bash
.venv/bin/repo-index init /path/to/repo --name my-repo
```
`init` is shorthand for `add` (register in Postgres) + `sync` (chunk, embed,
and index the current HEAD) + `install-hook` (wire up the git hook), in one
step. Individual steps can also be run on their own, e.g. to re-sync
on demand:
```bash
.venv/bin/repo-index sync my-repo
.venv/bin/repo-index status
```
`status` lists every registered repo with its path and last-synced commit.
## Staying current: the git hook
`install-hook` (also run by `init`) drops a `post-commit` hook into the
target repo's `.git/hooks/`. After every commit, it launches
`repo-index sync <name>` in the background (`nohup ... &`), logging to
`.git/repo-index-sync.log` inside the target repo, so commits are never
blocked or slowed down by re-indexing.
## Registering with Claude Code
The MCP server runs as a stdio process out of this project's venv:
```bash
claude mcp add repo-semantic-search -s user -- \
/Users/shlomi.hassan/projects/repo-semantic-search/.venv/bin/python -m repo_index.mcp_server
claude mcp list # should show repo-semantic-search - ✔ Connected
```
Registered at user scope, so `semantic_search` and `list_indexed_repos` are
available as tools in every Claude Code session (after a restart — newly
registered MCP servers only appear in *new* sessions). This coexists with
any other MCP servers already registered (e.g. an earlier, unrelated
`cocoindex-code` server from the exploratory phase); nothing here depends on
or conflicts with it.
## Verified working (2026-08-04)
Registered the MCP server (`claude mcp list` shows `repo-semantic-search -
✔ Connected`), then ran the full pipeline end-to-end against a real repo,
`~/projects/go-ip2country`:
- `repo-index init` registered the repo, indexed it (134 chunks across the
repo's Go source, tests, docs, and README), and installed the hook.
- Made a real commit in `go-ip2country`; the `post-commit` hook fired,
`repo-index-sync.log` showed a successful sync with no traceback, and
`repo-index status` picked up the new commit sha automatically.
- Ran `semantic_search` (via an in-memory MCP client) for `"how does the
rate limiter work"` scoped to `go-ip2country`: the top-ranked result
(score 0.80) was the README's "How the rate limiter works" section,
followed by the section on mutex locking/eviction — genuinely relevant,
correctly ranked results.
### Why `semantic_search` sets `ivfflat.probes` explicitly
The `code_chunks` table has a single `ivfflat` vector index shared across
all repos, and `semantic_search`'s `repo`-scoped queries filter with `WHERE
repo_name = $1` *after* the approximate-nearest-neighbor index scan. With
pgvector's default `ivfflat.probes = 1`, this could silently return fewer
than `top_k` results for a given repo even when more relevant matches exist
in the table — reproduced directly against Postgres: a query with
`top_k=5` returned only 2 rows through the ivfflat index at the default
probe count, but all 5 (including the actual
`internal/ratelimit/fixedwindow.go` implementation) with either a forced
sequential scan or `ivfflat.probes` raised to 10.
`semantic_search` now runs each query inside a transaction with `SET LOCAL
ivfflat.probes = 10`, which restored full recall in re-testing (see below).
This is a scoped, low-risk mitigation (session/transaction-local, no schema
change); a per-repo partial index or an HNSW index remain possible future
upgrades if recall issues resurface at larger scale, but aren't needed now.
## CLI reference
```
repo-index add <path> [--name NAME] # register a repo
repo-index sync <name> # chunk, embed, index current HEAD
repo-index install-hook <name> # install the post-commit hook
repo-index init <path> [--name NAME] # add + sync + install-hook
repo-index status # list registered repos + last sync
```
This server cannot be deployed
Maintenance
ActivitySlowing
ResponsivenessNo issues