Skip to main content
Glama
tgt-technology

code-index-pg

README.md
# code-index-pg

Unified code index for Hermes: **exact + semantic + call graph** search over the
TGT One monorepo, persisted in postgres. Replaces the code-index fork,
cocoindex and CodeGraphContext with one lightweight MCP server.

## Stack

- **postgres-local** (port 5432, DB `codeindex`) + pgvector 0.8.3
- **Ollama** `bge-m3` (1024 dims) for embeddings — no torch in this venv
- **tree-sitter-language-pack** (21 languages) for symbol extraction
- **fastmcp** — single MCP process, no extra daemons

## Install

```bash
git clone https://github.com/tgt-technology/code-index-pg.git
cd code-index-pg
uv sync
uv tool install --editable .
```

DB: la infraestructura (postgres + pgvector, ollama + `bge-m3`, base `codeindex`) se
levanta con `deploy/` — ver [Deploy](#deploy-entorno-desde-cero). El esquema de tablas lo
aplica el server al arrancar.

## Deploy (entorno desde cero)

El índice vive en postgres, no en archivos: reproducir el entorno es levantar la infra, no
copiar datos. Todo está en `deploy/`.

**Windows** (Docker Desktop + `uv`):

```powershell
powershell -NoProfile -File deploy\setup.ps1        # -Gpu si hay GPU NVIDIA
```

**Linux / WSL:**

```bash
deploy/setup.sh          # --gpu si hay GPU NVIDIA
```

El script levanta postgres (pgvector) + ollama, crea la base `codeindex` con la extensión
`vector`, descarga `bge-m3`, instala el server como uv tool y verifica cada punto al final
(falla con mensaje claro si algo quedó a medias).

Después, una sola vez, el índice inicial (~30 min para ~7K archivos):

```bash
uv run python -c "from code_index_pg.indexer_pipeline import index_project; print(index_project('/ruta/a/tu/codigo'))"
```

Registro en Hermes — el wrapper espera a que postgres acepte conexiones antes de lanzar el
server (sin eso, el server muere si la BD no está lista):

```bash
# Linux / WSL
hermes mcp add codeindex --command /ruta/al/repo/deploy/mcp-wrapper.sh

# Windows
hermes mcp add codeindex --command powershell --args -NoProfile -File C:\ruta\al\repo\deploy\mcp-wrapper.ps1
```

Solo los servicios van en Docker: el server MCP es stdio y corre en el host, usando los
defaults del código (`127.0.0.1:5432/codeindex`, `http://localhost:11434`).

## Run

```bash
# MCP server (stdio — used by the proxy wrapper)
code-index-pg --project-path ~/develop

# Full index (one-time, ~10 min for ~7K files)
python -m code_index_pg.indexer_pipeline  # via test harness, or:
uv run python -c "from code_index_pg.indexer_pipeline import index_project; print(index_project('/home/jam/develop'))"
```

## MCP tools

| Tool | Description |
| --- | --- |
| `set_project_path` | Index a project and start the incremental watcher |
| `update_index` | Incremental re-index: add missing files, drop deleted ones (no re-embed) |
| `search_code` | Literal/regex search over symbol names + file paths (paginated) |
| `search_semantic` | Embedding similarity search (bge-m3, pgvector cosine) |
| `callers_of` | Symbols that call the given symbol |
| `callees_of` | Symbols called by the given symbol |
| `impact_analysis` | 2-level recursive callers + affected files |
| `find_files` | Glob search over indexed files |
| `get_symbol_body` | Signature, docstring, location of a symbol |
| `get_file_summary` | Language, line count, symbols of a file |
| `get_status` | Indexed projects + counts + active project |
| `remove_project` | Remove an indexed project and all its data in cascade (rejects the active project) |

## Hermes integration

- Proxy: port 3104 in `~/.hermes/scripts/mcp-proxies.sh` (supervisor systemd)
- Wrapper: `~/.hermes/scripts/mcp-codeindex-wrapper.sh`
- Register: `hermes mcp add codeindex --url http://127.0.0.1:3104/mcp`

## Environment

| Var | Default | Notes |
| --- | --- | --- |
| `CODEINDEX_DATABASE_URL` | `postgres://postgres:postgres@127.0.0.1:5432/codeindex` | local tool |
| `OLLAMA_URL` | `http://localhost:11434` | embeddings |
| `CODEINDEX_EMBED_MODEL` | `bge-m3` | must match pgvector dims (1024) |
| `CODEINDEX_PROJECT` | unset | active project (set via `set_project_path`) |

## Tests

```bash
uv run pytest tests/ -q   # unit + contract (in-memory MCP client)
```

## Spec / OpenSpec

Living spec: `openspec/specs/code-index/spec.md` — archived changes in
`openspec/changes/archive/` (`unified-code-index`, `project-directory-management`).

## License

MIT — free to use, modify and redistribute. See [LICENSE](LICENSE).

TDQS

A3.6/5.0

Scored across 12 tools

Disambiguation5/5

Each tool targets a distinct operation: indexing, searching (literal vs semantic), graph queries (callers/callees), file operations, and project management. No two tools appear to overlap in purpose.

Naming Consistency4/5

Most tool names follow a clear verb_noun pattern (set_project_path, search_code, get_status). Minor exception: 'callers_of' and 'callees_of' use a different prepositional style, but still readable and predictable.

Tool Count5/5

12 tools is well within the ideal range for a code indexing/search server. Each tool serves a distinct function needed for code navigation and project management, with no redundancy.

Completeness5/5

The surface covers the full lifecycle: project setup, incremental updates, multiple search modes, symbol graph queries, file and symbol detail, and project removal. Missing features like batch re-index or cross-project search are minor and not essential.

Maintenance

ActivityMaintained
ResponsivenessNo issues