cindex
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@cindexsearch for where we handle file uploads"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
cindex
Local, offline semantic code search: tree-sitter AST chunks → local GGUF embeddings (Qwen3-Embedding-0.6B via llama-server) → SQLite → brute-force vector search → CLI + MCP tool for coding agents. No network calls, no API keys. The index is content-addressed, so unchanged code is never re-embedded — reformatting a repo or moving files/functions costs zero inference.
What gets indexed: Python, C++, JavaScript (AST chunks: functions, class skeletons with bodies stripped, merged top-level blocks), Markdown (heading sections), plain text (100-line windows). Other file types are currently skipped (see roadmap).
Setup
macOS
git clone <this-repo> && cd code-indexer
uv sync # install python deps from the lockfile
brew install llama.cpp # provides llama-serverLinux
git clone <this-repo> && cd code-indexer
curl -LsSf https://astral.sh/uv/install.sh | sh # if uv is not installed
uv sync
# llama.cpp: use a prebuilt release from https://github.com/ggml-org/llama.cpp/releases
# or build from source:
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j
sudo cp build/bin/llama-server /usr/local/bin/ # or add build/bin to PATH
cd ..Model weights (both platforms, one time, ~630 MB, gitignored)
uv tool install "huggingface_hub[cli]"
hf download Qwen/Qwen3-Embedding-0.6B-GGUF Qwen3-Embedding-0.6B-Q8_0.gguf --local-dir models/Related MCP server: CodeGrok MCP
Starting it
Terminal 1 — the inference server. Leave it running; Ctrl-C stops it.
uv run cindex serveAlways start llama-server through cindex serve: it pins flags the client depends
on (--pooling last -ub 4096 -b 4096; without -ub the server crashes on long
inputs — details in docs/llama-server-pinned.md).
Terminal 2 — build the index, then query.
uv run cindex init # create the database (one time)
uv run cindex index # embed the repo; re-run any time, only changes cost inference
uv run cindex query "where do we decide if a file needs re-embedding" -k 5Every command fails fast with the fix in the message (server down, no index yet, model/config mismatch), so if something's wrong the output says what to run.
Using it
uv run cindex query "natural language question" # top 8 by meaning
uv run cindex query "..." -k 3 # fewer results
uv run cindex query "..." --path src/ # restrict to a subtree
uv run cindex query "..." --no-instruct # raw query embedding (A/B)
uv run cindex index # refresh after editing codeResults are score file:start-end [chunk-type] symbol plus the live snippet read
from disk. Stale results self-heal: if a file changed since indexing, the hit is
re-indexed inline and correct line numbers are returned.
Multiple indexes (one config = one root = one database)
The default config.toml indexes this repo. To index any other tree, copy
documents.toml's pattern — set root (relative to the config file) and a
distinct db — and pass --config:
uv run cindex --config documents.toml index # workspace-wide index
uv run cindex --config documents.toml query "..." --path 6106 # scope to one projectQuerying the default config only searches this repo; if you expected results from
a sibling project, you queried the wrong index — add --config.
Agents (MCP)
Opening this repo in Claude Code auto-registers the search_code tool via the
committed .mcp.json (approve the server when prompted). AGENTS.md instructs
agents to prefer it over grep for meaning-based lookup. Requirements: cindex serve
running; the MCP server creates/updates the index itself at session start and
repairs stale files lazily at query time.
To give agents in ANY directory a workspace-wide index, register at user scope:
claude mcp add cindex --scope user -- uv --directory /abs/path/to/code-indexer run cindex --config documents.toml mcpAgents can pass path_prefix to search_code to stay inside one subproject.
Upcoming improvements
More languages — TypeScript, Go, Rust, Java: each is one tree-sitter wheel + one walker extension entry + one chunker
LangSpec.Catch-all indexing — unknown text file types as blob windows so nothing is invisible, just coarser.
Chunker versioning — stamp
chunker_versionin meta and force re-chunk on upgrade (today an unchanged file keeps its old chunking).Finer prose chunking — paragraph windows for
.txtinstead of 100-line blobs.ANN search (sqlite-vec) once brute-force matmul exceeds ~50 ms (~10⁶ chunks).
int8 quantization — the
encodingcolumn is already reserved for it.Watcher daemon — filesystem events instead of explicit
cindex index.Hybrid ranking — path/symbol signal fused at rank time (never into content vectors), plus optional reranker.
Bench harnesses — speed + recall regression tracking with tagged baselines.
GPU offload flags — config-only change when needed.
Layout
src/cindex/ config db walker hasher chunker embedder indexer search resolver cli mcp_server
docs/ llama-server-pinned.md — the pinned inference server contract
config.toml default index (this repo) · documents.toml — example second index
.mcp.json wires Claude Code to `cindex mcp` · AGENTS.md — usage instruction for agentsMaintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Tools
Related MCP Servers
- Flicense-qualityDmaintenanceProvides semantic code search capabilities that run 100% locally using EmbeddingGemma embeddings. Enables finding code by meaning across 15 file extensions and 9+ programming languages without API costs or sending code to the cloud.237
- Alicense-qualityFmaintenanceEnables semantic code search for AI assistants by indexing codebases with embeddings and Tree-sitter, returning relevant snippets via natural language queries.15MIT
- AlicenseAqualityDmaintenanceProvides semantic code search over codebases using local embeddings with natural language queries. Supports hybrid search, file watching, and respects .gitignore.115MIT
- Flicense-qualityDmaintenanceEnables local semantic code search across repositories using natural language, with AST-aware chunking and hybrid vector/FTS5 retrieval.
Related MCP Connectors
Token-efficient search for coding agents over public and private documentation.
Persistent memory and knowledge management for AI agents with semantic search and 50+ tools.
Securely search and manage workspace context files for AI agents and teams.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/aarontrmartin/code-indexer'
If you have feedback or need assistance with the MCP directory API, please join our Discord server