anyindex-mcp
anyindex-mcp is a fully local MCP server that indexes a codebase and answers semantic/keyword code-search queries over it — no API keys, no telemetry, no external calls.
anyindex_search— Semantic, keyword, or hybrid (RRF-fused) search over the entire indexed repo; returns file, line range, entity, distance, and confidence.topKcounts files, max 25.find_references— Lexical lookup of where an identifier is declared and used (explicitly not a call graph); optional file scoping andmaxResults.get_file_outline— Lists functions, classes, and methods in one file with line ranges and signatures; cheaper than reading the file.index_status— Reports readiness,degradedstate, file/chunk counts, staleness, breakdown (AST vs fallback vs non-code), and live indexing progress.index_update— Incremental reindex + embed of changed files, reusing chunks by content hash; runs in the background.index_rebuild— Destructive full rebuild from scratch, recomputing every embedding; runs in the background.ping— Liveness probe returning the resolved config (root, model, node version, index and model paths) to verify client wiring.
Additional traits: AST-aware chunking via tree-sitter, offline ONNX embeddings (Jina v2 base code), one SQLite + FTS5 file per project, respects git ignore rules, and supports --watch for keeping the index fresh.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@anyindex-mcpfind where user login is handled in my codebase"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
anyindex-mcp
Fully local MCP server for indexing a codebase and answering semantic search queries about it. No external API calls, no telemetry, no server process to run.
AST-aware chunking via tree-sitter — chunks follow function and class boundaries, never cut mid-statement
Offline embeddings via ONNX in-process; after the first download the model works with no network
Hybrid search — vector + BM25 fused with RRF
One SQLite file in the project, plus an FTS5 index
Works on Windows without Visual Studio Build Tools — see Platform
Quickstart
Wire it into a client — nothing to install, see Client configuration:
npx --package anyindex-mcp anyindex-mcp-serverBuild the index once from your project, so the first search has something to read:
npx --package anyindex-mcp anyindex-mcp reindex --root /path/to/project
npx --package anyindex-mcp anyindex-mcp search --root /path/to/project "user authentication flow"After a global npm install -g anyindex-mcp the npx --package … prefix drops away
and the commands are just anyindex-mcp reindex and anyindex-mcp search.
A second reindex only touches changed files. To keep the index fresh while you
edit, pass --watch to the server, not to the CLI:
npx --package anyindex-mcp anyindex-mcp-server --watchTo work on anyindex-mcp itself, from a checkout:
npm install
npm run build
npm run probe # environment check, prints ms/chunkRelated MCP server: semantic-search-mcp
Requirements
Node ^22.19.0 || >=24.0.0. Node 23.x does not qualify at any patch level.
Embedding dominates the wall clock: plan against ~1 s/chunk on 4 cores, not the optimistic probe figure. If a run stalls, it is memory pressure — free RAM and retry. Measured numbers in docs/06-environment.md §14.
Install
To use it as an MCP server, install nothing. Every client below runs the
package through npx, which downloads it on first start and caches it:
npx --package anyindex-mcp anyindex-mcp-serverTo use the command line from any project, either prefix each call the same way or install it once:
npm install -g anyindex-mcp
anyindex-mcp --versionnpm install -g needs elevated rights on some Windows setups. The npx form works
everywhere and is otherwise identical.
To work on anyindex-mcp itself, from a checkout:
npm install
npm run build
npm run probe # environment check, prints ms/chunkprobe checks Node version, sqlite-vec, code-chunk, the embedding model, and
write permissions, and prints throughput numbers. Exit code 1 means a blocking
failure — the stack should be reconsidered before writing code.
The package also installs as a dependency and puts both executables on your path. Checked from a tarball into an empty project, with npm's install scripts blocked:
npm install ./anyindex-mcp-1.0.1.tgz
npx anyindex-mcp helpUsage
Search output looks like this:
1. src/auth/login.ts:14-24 loginUser distance 0.957 (moderate)
export async function loginUser(email: string, password: string) {
const row = await db.query('select * from users where email = ?', [email])
...Running as an MCP server
The server speaks MCP over stdio. Nothing is written to stdout except JSON-RPC framing; all diagnostics go to stderr, so the channel stays clean. It takes no arguments and needs no configuration file.
npx --package anyindex-mcp anyindex-mcp-serverWire it into a client as described below.
Command-line reference
anyindex-mcp reindex --root /path/to/project # build the index
anyindex-mcp update --root /path/to/project # incremental
anyindex-mcp status --root /path/to/project # state, --json for machine output
anyindex-mcp search --root /path/to/project "user authentication flow"
anyindex-mcp probe # environment and hardware
anyindex-mcp benchmark # quality benchmark
anyindex-mcp --version # package version
anyindex-mcp help # same as --helpreindex and update exit 1 if --root does not exist. Without that check a
mistyped path scans nothing and reports a finished reindex.
search supports --mode semantic|keyword|hybrid, --topK, and --json; --help prints usage. Exit code is 1 when the index is not ready, so shell scripts can branch on it. Search refuses to run against an unready index rather than returning plausible-looking results.
topK counts files, not chunks: one file never takes two slots, so you may get
fewer results than you asked for.
Client configuration
Every client runs the same command and differs only in the config file it reads.
All of them run the package through npx, so there is nothing to install first:
command: npx
args: -y --package anyindex-mcp anyindex-mcp-server
env: ANYINDEX_LOG_LEVEL=warn (optional)--package is load-bearing. npx resolves a package name, not a binary name,
and this package ships two binaries — anyindex-mcp (the CLI) and
anyindex-mcp-server (the MCP server). Writing npx -y anyindex-mcp-server
resolves nothing:
npm error could not determine executable to runclaude mcp add writes the entry for you. Everything after -- goes to the
server untouched, which is what keeps -y from being read as Claude Code's own
flag:
claude mcp add anyindex-mcp -- npx -y --package anyindex-mcp anyindex-mcp-serverThe same entry by hand in .mcp.json at the project root:
{
"mcpServers": {
"anyindex-mcp": {
"type": "stdio",
"command": "npx",
"args": ["-y", "--package", "anyindex-mcp", "anyindex-mcp-server"],
"env": { "ANYINDEX_LOG_LEVEL": "warn" }
}
}
}~/.cursor/mcp.json for every project, .cursor/mcp.json for one, or
Settings → Customize → MCP to add it through the UI.
{
"mcpServers": {
"anyindex-mcp": {
"type": "stdio",
"command": "npx",
"args": ["-y", "--package", "anyindex-mcp", "anyindex-mcp-server"]
}
}
}.vscode/mcp.json in a project, or the user-profile mcp.json from the
MCP: Open User Configuration command — servers go under a top-level
servers object there:
{
"servers": {
"anyindex-mcp": {
"type": "stdio",
"command": "npx",
"args": ["-y", "--package", "anyindex-mcp", "anyindex-mcp-server"]
}
}
}The portable alternative is .mcp.json at the project root, which other clients
read too, under mcpServers:
{
"mcpServers": {
"anyindex-mcp": {
"type": "stdio",
"command": "npx",
"args": ["-y", "--package", "anyindex-mcp", "anyindex-mcp-server"]
}
}
}Claude → Settings → Developer → Edit Config.
macOS:
~/Library/Application Support/Claude/claude_desktop_config.jsonWindows:
%APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"anyindex-mcp": {
"command": "npx",
"args": ["-y", "--package", "anyindex-mcp", "anyindex-mcp-server"]
}
}
}Claude Desktop has no working directory of its own, so it launches the server from
its own config directory. See which folder gets indexed
below — for Desktop, --root is the one thing worth setting.
Quit and restart Claude Desktop: it reads the file at launch, not on change.
Settings → AI → MCP Servers → Add Local Server, or open the settings file
directly (zed: open settings file). Zed's key is context_servers:
{
"context_servers": {
"anyindex-mcp": {
"command": "npx",
"args": ["-y", "--package", "anyindex-mcp", "anyindex-mcp-server"],
"env": {}
}
}
}mcp.<name>.command is an array here, not a string, and there is a separate
cwd. The schema is validated at startup and rejects unknown fields.
Project config: ./opencode.json, ./opencode.jsonc or .opencode/opencode.json
at the repository root. The configuration is read once at startup and is not
reloaded — restart opencode after an edit.
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"anyindex-mcp": {
"type": "local",
"command": ["npx", "-y", "--package", "anyindex-mcp", "anyindex-mcp-server"],
"cwd": ".",
"environment": { "ANYINDEX_LOG_LEVEL": "warn" },
"enabled": true,
"timeout": 120000
}
}
}timeout matters: the first anyindex_search loads the model into a worker and
took 2745 ms on a warm index, against 145 ms for the next call. opencode defaults
to 5000 ms, which is enough — a client with a shorter timeout drops the first call
while everything is actually working.
Almost every MCP client reads a stdio server the same way:
command: npx
args: -y --package anyindex-mcp anyindex-mcp-server
env: ANYINDEX_ROOT=/absolute/path (optional, see below)On Windows, use npx.cmd if your client cannot launch npx through a shell. It
takes no arguments, needs no API key, and writes its log to stderr, so stdout
carries nothing but JSON-RPC.
To confirm a client started it correctly, call ping — it reports the resolved
configuration rather than just answering. See
docs/10-clients.md
for per-client notes and what has and has not been verified.
Which folder gets indexed
No project path is required in the config. The working folder is taken from
wherever the agent was started, resolved up to the nearest enclosing repository
root. Start the agent in ~/work/myrepo/src/api and it indexes ~/work/myrepo,
not the subdirectory. Outside a repository it uses the working directory as it
stands.
That also means the server has to be started from the project. A client that
spawns MCP servers from its own config directory rather than the project would
resolve the wrong root; ping reports the root it actually chose, so this is
one call to check:
{ "echo": "pong", "root": "/home/you/work/myrepo", "model": "jinaai/jina-embeddings-v2-base-code",
"node": "v24.21.0", "index": "/home/you/work/myrepo/.anyindex/index.db",
"models": "/home/you/.cache/anyindex-mcp/models" }To pin the folder regardless of where the client starts from, set ANYINDEX_ROOT
in env to an absolute path.
Before the first search
A fresh install has no index. Until one is built, anyindex_search refuses to
answer rather than returning plausible-looking results, and index_status says so.
Build it once, from the project:
npx --package anyindex-mcp anyindex-mcp reindex --root .After that a second reindex only touches changed files.
Configuration
All settings are optional. The server reads flags first, then environment variables, then defaults.
Flag | Environment variable | Default |
|
| nearest repository root above the working directory, else the working directory |
|
| nearest repository root above |
|
|
|
|
| user cache: |
|
|
|
|
|
|
|
|
|
|
|
|
— |
|
|
— |
|
|
Chunking and batching internals (chunkSizeBytes, maxFileBytes, batchSize) are fixed in src/config.ts and intentionally not exposed as flags or environment variables.
Only the index goes into the project: .anyindex/ holds the database and nothing
else. The model is ~157 MB, identical for every project on the machine, so it is
downloaded once into the user cache — <root> stays clean, and deleting
node_modules does not throw it away. Point --models elsewhere to share one
cache between accounts or to keep it on a different volume.
No API keys. The embedding model is Apache-2.0 and ungated, so nothing secret is ever configured — which also means there are no *_TOKEN-style variables for MCP clients to strip from subprocess environments.
The model is fetched from Hugging Face on first use, not by npm install, and lands in the user cache. After that first download nothing reaches the network, so --offline works from a warm cache and fails with a clear error from a cold one.
Ignored files
Ignore rules follow git's own precedence, verified against git check-ignore:
Source | Scope | Precedence |
| repository | lowest |
| directory containing the file | per directory |
|
| highest |
Nested ignore files are honoured: src/.gitignore applies to src/, and a deeper
file overrides a shallower one, including re-including a path with !. node_modules/
and .git/ are excluded programmatically and cannot be re-enabled by any rule.
Paths are relative to --root.
Tools
Tool | Annotations | Description |
| read-only | Semantic + keyword search over the whole index. Returns file, lines, entity, and the match distance |
| read-only | Where an identifier is declared and used. Lexical matches, not a call graph |
| read-only | Entities defined in one file, with line ranges. Cheaper than reading the file |
| read-only | Readiness, counts, staleness, and progress of a running index job |
| writes | Incremental reindex and embed. Runs in the background |
| destructive | Full rebuild, recomputing every embedding |
| read-only | Liveness probe returning the resolved configuration |
Parameters:
Tool | Parameters |
|
|
|
|
|
|
| none |
| none |
| none |
| none |
index_update and index_rebuild return immediately and work in the background, so
the client's request timeout does not limit indexing. Poll index_status until
running is false.
The timeout does matter for anyindex_search: the first call loads the model and
takes seconds. Measured 2745 ms on a warm model, then 145 ms and 119 ms.
anyindex_search refuses to run when index_status reports the index unready or
degraded, rather than returning weaker results that look normal. A half-vectorized
index counts as degraded: index_rebuild is the fix. Per-client wiring and timeouts
are in the client notes; see ADR-010 for why the refusal
exists.
Measured quality
The indexer, hybrid search and watcher work end to end. 118 tests pass. Verified
on Linux x64, Windows x64 and macOS x64 — CI runs the build, the type check and
the test suite on all three, on Node 22.19 and 24, plus probe on each.
Retrieval quality was measured on 520 questions mined from commit history in Flask,
Gson and Fastify, against a grep baseline charged with the same tokenizer: the
correct file reaches the top 3 in 0.531 of cases versus 0.110 for grep, at
1080 context tokens per question versus 3800.
Two limits are measured and unfixed:
Russian queries retrieve roughly 0.2 further than English ones. A multilingual model scores the same, so this is a property of matching Russian against English code, not of the model.
No relevance threshold exists — across three models the distance gap between the worst positive and best negative case is negative. Results carry a distance and a confidence label instead of being silently filtered.
Quality also falls as the corpus grows, which is currently the largest measured limit: top-3 is 0.711 on Flask, 0.473 on Gson, 0.328 on Fastify. Benchmark, per-language numbers and the failure analysis are in docs/06-environment.md §18.
Platform
Target platform is Windows (ADR-011). Rules that follow from it:
Native modules ship a prebuilt binary inside the package; no Visual Studio Build Tools are needed
Paths are normalized to POSIX separators before storage, so an index built on Windows opens on Linux and back
A
.cmdwrapper is installed alongside the POSIX entry pointchokidarreports paths with forward slashes whilerootarrives with backslashes, so anything comparing them must go throughpath.relative. A plain string comparison makes the watch root look like a path outside itself, and the watcher then reports nothing without any error
npm 11.21 or newer is required, and engines says so. This is not about a
script we need: npm below 11.21 does not block dependency install scripts, and it
synthesises node-gyp rebuild for any package that ships a binding.gyp without
an install script — better-sqlite3 is exactly that case. The rebuild needs MSVC,
the Windows runner does not have it, and the node-gyp bundled with npm 10 cannot
even recognise the Visual Studio 18 that is installed there:
gyp ERR! find VS unknown version "undefined" found at "C:\Program Files\Microsoft Visual Studio\18\Enterprise"Nothing here needs a script: onnxruntime-node, better-sqlite3 and the six
tree-sitter-* grammars all carry their binaries in the tarball, so with scripts
blocked they load without npm install-scripts approve. This was found by CI, not
by inspection — see Measured quality.
Development
npm run build # compile TypeScript and prepare bin entries
npm run typecheck # type check without emitting
npm test # integration tests against a real stdio server
npm run probe # environment verificationExternal quality benchmark. Mines questions from commit history in Flask, Gson and Fastify at pinned commits and compares against a grep baseline. Clones the three repositories (~46 MB) and needs several hours for the first index of each; later runs reuse the index and take minutes.
node scripts/bench-external.mjs --corpus flask # also: gson, fastifyThe scripts behind the analysis in docs/06-environment.md §18 live next to it:
rank-sweep.mjs compares ranking strategies on built indexes, doc-regression.mjs
checks that documentation queries keep working, language-diagnosis.mjs measures
how answerable the questions are, and scale-control.mjs re-measures the same
questions on a smaller index.
Documentation
These are the design notes behind the release — measurements, decisions and the open items. They are not shipped in the npm package, so the links point at the repository.
File | Contents |
Index and decisions summary | |
Verified facts about every dependency, with sources | |
Components, schema, tool design | |
Implementation stages with acceptance criteria | |
Risk register and decision log (17 ADR) | |
Survey of existing tools and why they were rejected | |
Environment measurements, per platform | |
Bun/Node comparison, measured | |
Why there is no standalone binary | |
Inventory of what is still missing | |
Per-client configuration and timeouts |
License
MIT
Available Tools
7 toolsanyindex_searchARead-onlyIdempotent
Semantic and keyword search over the ENTIRE locally indexed codebase. Every indexed file is already available to you — do not ask the user to paste code, and do not read files directory-by-directory looking for it. Use this when the answer is not in your current context, when you need to locate a function, class or feature by name or by meaning, or when you need to understand how parts of the project fit together. Skip it when the answer is already in context, or for general programming questions unrelated to this repository. Each result carries file path and line numbers. If no results come back, the code may not exist here — say so instead of inventing it.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | hybrid | |
| topK | No | Maximum number of results. Counts files, not chunks: you get at most one chunk per file, so fewer results may come back. Use get_file_outline for the rest of a file. | |
| query | Yes | Natural language query, or an exact symbol name |
Output Schema
| Name | Required | Description |
|---|---|---|
| hits | Yes | |
| index | Yes | |
| total | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered. The description adds behavior beyond that: results carry file path and line numbers, the tool counts files not chunks, and it prescribes what to do when no results return (say so rather than invent). It does not describe pagination or ranking behavior, so it stops short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with scope and the most important behavioral instruction (don't ask for pasted code, don't crawl directories). Every sentence carries actionable content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, yet the description still notes path/line payload and the zero-result case. For a search tool with annotations covering safety, nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67% — topK and query are documented in the schema, but mode has only an enum with no explanation of hybrid vs semantic vs keyword. The description text offers no parameter guidance at all, so it neither compensates for the gap nor adds meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('search') and resource ('the ENTIRE locally indexed codebase') and names both modalities (semantic and keyword). An agent can immediately tell this apart from siblings like index_status or get_file_outline.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly enumerates when to use ('answer not in current context', 'locate a function/class/feature', 'understand how parts fit together') and when to skip ('answer already in context', 'general programming questions unrelated to this repository'). It also names a follow-up alternative, get_file_outline, and warns against directory-by-directory reading.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_referencesARead-onlyIdempotent
Finds where an identifier is defined and where it is used, as file and line numbers. Use it before editing a shared function or type, to see what depends on it, or when you know the name but not the file. These are LEXICAL matches: a same-named identifier in another class, a comment or a string counts as a match too. It is not a call graph — it does not resolve which Charge a given call site means. Prefer anyindex_search when you want code by meaning rather than by name.
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | Restrict to one file, path relative to the indexed root | |
| symbol | Yes | Exact identifier to look for, e.g. startWatcher | |
| maxResults | No | ||
| includeDefinitions | No | Include the sites where the identifier is declared |
Output Schema
| Name | Required | Description |
|---|---|---|
| symbol | Yes | |
| truncated | Yes | |
| otherFiles | Yes | |
| references | Yes | |
| definitions | Yes | |
| lexicalOnly | Yes | |
| matchedChunks | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description adds substantive semantics beyond them: matches are LEXICAL, so same-named identifiers in other classes, comments and strings count, and call sites are not resolved. This is exactly the caveat an agent needs to trust results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, zero filler, ordered from what it does to when to use it to the lexical caveat to the alternative. Every sentence carries distinct decision-relevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return format need not be restated. With annotations covering the safety profile and the description covering lexical-matching limits and sibling routing, nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, with file, symbol and includeDefinitions documented in the schema itself. The description clarifies the matching semantics of `symbol` (exact identifier) and implies definition-site results, but adds no syntax or default details for file or maxResults. Baseline 3 fits when the schema does most of the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('finds where an identifier is defined and where it is used') plus the output shape (file and line numbers). It explicitly separates itself from the sibling anyindex_search, so an agent can disambiguate without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete triggers ('before editing a shared function or type', 'when you know the name but not the file') and an explicit exclusion rule ('It is not a call graph'), plus a named alternative with its selecting condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_file_outlineARead-onlyIdempotent
Lists the symbols defined in one file — functions, classes, methods — with their line ranges and signatures. Cheaper than reading the file when you only need to know what it defines, or need to pick a line range to read. The index may be stale for very recent edits; verify with Read when exactness matters.
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | Path relative to the indexed project root, POSIX separators |
Output Schema
| Name | Required | Description |
|---|---|---|
| file | Yes | |
| note | No | |
| indexed | Yes | |
| entities | Yes | |
| language | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds genuinely useful non-annotation context: the index may be stale for very recent edits, and Read should be used to verify exactness. That staleness/verification caveat is the kind of trait structured fields cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the output shape, then the usage tradeoff, then the staleness caveat. Each sentence carries distinct information and none is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return format need not be explained, and annotations cover the safety profile. Combined with the usage rule and staleness caveat, everything an agent needs to call this correctly is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema coverage, so the schema already defines 'file' as a project-root-relative POSIX path. The description adds no syntax or formatting detail beyond what the schema states, so the baseline of 3 for high-coverage schemas applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb+resource+scope: 'Lists the symbols defined in one file — functions, classes, methods — with their line ranges and signatures.' An agent immediately knows both the operation and the payload. It does not explicitly distinguish itself from siblings like find_references or anyindex_search, which is the only gap keeping this from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit decision rule ('Cheaper than reading the file when you only need to know what it defines, or need to pick a line range to read') and names the alternative tool (Read) for the case where exactness matters. When-to-use and when-to-verify are both covered with no inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
index_rebuildADestructive
Discards the whole index and rebuilds it from scratch, including recomputing every embedding. Slow — minutes on a large repository. Use only when index_status reports an incompatibility, or after changing the embedding model.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| files | Yes | |
| phase | Yes | |
| ready | Yes | |
| stale | Yes | |
| chunks | Yes | |
| reason | No | |
| reused | Yes | |
| running | Yes | |
| skipped | Yes | |
| degraded | Yes | |
| embedded | Yes | |
| progress | Yes | |
| astChunks | Yes | |
| vecVersion | Yes | |
| lastUpdated | Yes | |
| nonCodeChunks | Yes | |
| fallbackChunks | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (destructiveHint=true), and the description adds substantive context beyond them: that the entire index is discarded, every embedding is recomputed, and the operation takes minutes on large repositories. The cost/time warning is exactly the kind of behavioral detail an agent needs before invoking a destructive, expensive tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: the operation, its consequence/cost, then the gating condition. Front-loaded with the destructive action, no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the description covers the remaining gaps for a zero-param destructive tool: scope of destruction, performance cost, and preconditions. Nothing an agent needs to decide whether to call it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description correctly introduces no parameter semantics because there are none to describe.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Discards the whole index and rebuilds it from scratch') with the distinguishing scope word 'whole' that separates it from an incremental update. An agent can immediately tell this is the full-rebuild path versus the refresh path.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rule with two concrete triggers: 'only when index_status reports an incompatibility, or after changing the embedding model.' It even names the sibling tool (index_status) that gates the decision, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
index_statusARead-onlyIdempotent
Reports whether the codebase index is ready to search, how many files and chunks it covers, and whether indexing is currently running. After calling index_update or index_rebuild, poll this to find out when the work finished. Also reports whether any subsystem failed, so you never rely on a silently broken index.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| files | Yes | |
| phase | Yes | |
| ready | Yes | |
| stale | Yes | |
| chunks | Yes | |
| reason | No | |
| reused | Yes | |
| running | Yes | |
| skipped | Yes | |
| degraded | Yes | |
| embedded | Yes | |
| progress | Yes | |
| astChunks | Yes | |
| vecVersion | Yes | |
| lastUpdated | Yes | |
| nonCodeChunks | Yes | |
| fallbackChunks | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered. The description adds useful behavioral context beyond that: it reports whether any subsystem failed, warning the agent not to rely on a silently broken index. It stops short of describing polling frequency or consistency guarantees, but it meaningfully extends the annotation profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place. The core purpose is front-loaded, the polling instruction follows naturally, and the failure-detection note is a useful final clause with no wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter status tool with a dedicated output schema, the description provides everything needed: what it reports, when to call it, and the failure-detection guarantee. Return-value details can be left to the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There are no parameter semantics to clarify, and the schema coverage is already 100% for the empty object.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it reports whether the index is ready, how many files/chunks it covers, whether indexing is running, and whether subsystems failed. It also names sibling tools (index_update, index_rebuild) in a way that distinguishes this status tool from the mutation tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is explicitly framed: poll this after calling index_update or index_rebuild to know when the work finished. The description leaves no ambiguity about when this tool is appropriate, and it names the specific triggering siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
index_updateAIdempotent
Indexes files that changed since the last run and embeds the new chunks. Safe to call after edits; existing chunks are reused by content hash. Use it when the repository has been modified and the index is stale.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| files | Yes | |
| phase | Yes | |
| ready | Yes | |
| stale | Yes | |
| chunks | Yes | |
| reason | No | |
| reused | Yes | |
| running | Yes | |
| skipped | Yes | |
| degraded | Yes | |
| embedded | Yes | |
| progress | Yes | |
| astChunks | Yes | |
| vecVersion | Yes | |
| lastUpdated | Yes | |
| nonCodeChunks | Yes | |
| fallbackChunks | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare idempotentHint=true, destructiveHint=false; the description goes further by explaining WHY it's safe: 'Safe to call after edits; existing chunks are reused by content hash.' This adds mechanism beyond the annotation, exactly what the lower bar rewards.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight clauses, front-loaded with the action, then the safety guarantee, then the usage condition. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No params, output schema exists so return values needn't be explained, and annotations cover the safety profile. The description adds the incremental/hash-reuse semantics and the trigger condition. Complete for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so the schema provides nothing and the baseline is 4. The description correctly implies no inputs are needed (it operates on repository state), and adds the 'since the last run' semantics that define how the tool behaves without arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Indexes files') plus the incremental scope ('that changed since the last run') and the secondary effect ('embeds the new chunks'). This distinguishes it from siblings like index_rebuild (full rebuild) and index_status (reporting).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use: 'Use it when the repository has been modified and the index is stale.' Sibling index_rebuild exists but is not contrasted; still, the trigger condition is stated clearly enough to route the agent, so this lands at 5 as clear context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pingARead-onlyIdempotent
Liveness probe. Returns the resolved configuration so client setup can be verified.
| Name | Required | Description | Default |
|---|---|---|---|
| echo | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| echo | Yes | |
| node | Yes | |
| root | Yes | |
| index | Yes | |
| model | Yes | |
| models | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds real value beyond them by disclosing that it returns the resolved configuration for client verification, which is the actual behavioral purpose of a ping in this context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the tool's identity and followed by the return behavior. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is unnecessary, and the description covers the tool's purpose and return intent well. The only shortfall is the undocumented 'echo' parameter, which for a trivial probe is a minor omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter 'echo' has 0% schema description coverage, so the description carries the full burden – yet it never mentions 'echo' or what it does. A caller cannot tell whether supplying echo changes the response, which is a meaningful gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource for a diagnostic tool: 'Liveness probe' plus 'Returns the resolved configuration.' An agent can distinguish this from the index-oriented siblings (index_status, anyindex_search, etc.) without opening the schema, though it never explicitly names those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage ('so client setup can be verified') but gives no explicit when-to-use or when-not-to-use guidance, and no alternatives. For a self-evident liveness probe this is adequate but not instructive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v1.0.0- First observed
anyindex_search - First observed
find_references - First observed
get_file_outline - First observed
index_rebuild - First observed
index_status - First observed
index_update - First observed
ping
TDQS
Scored across 7 tools
Tools cover distinct lifecycle stages (status, update, rebuild) and distinct retrieval modes (semantic search, symbol outline, lexical references). Minor overlap between anyindex_search and find_references when locating an identifier by name, but descriptions clarify the intended use cases.
Mixed conventions: index_* prefix for management tools is consistent, but anyindex_search has a redundant server prefix, and verbs like get_file_outline/find_references coexist with noun-like index_status and bare ping. Still readable but not uniform.
Seven tools are well-scoped for a codebase index/search server. Each tool has a clear role and no redundant or bloated surface.
Covers full index lifecycle (status, update, rebuild), search, file outline, references, and liveness. No obvious missing operation for the stated purpose.
Maintenance
Related MCP Connectors
Shared memory for coding agents. Stop re-explaining your codebase every session.
Code intelligence for coding agents: semantic, AST, graph, and full-text search. 279+ languages.
Project memory, semantic code search, and grounded agent context.
Code intelligence for LLMs. Analyze, search, and retrieve code from any public git repository.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceA local code indexing and search library that enables AI agents to perform semantic and keyword searches across codebases using tree-sitter and SQLite. It provides tools for indexing projects and finding precise code definitions without requiring external APIs, Docker, or server infrastructure.1,307 npm344MIT
- AlicenseAqualityDmaintenanceProvides semantic code search over codebases using local embeddings with natural language queries. Supports hybrid search, file watching, and respects .gitignore.115MIT
- FlicenseNot gradedqualityDmaintenanceEnables local semantic code search across repositories using natural language, with AST-aware chunking and hybrid vector/FTS5 retrieval.-
- AlicenseAqualityAmaintenanceSemantic codebase search + persistent working memory for AI code editors. Local, zero-config, MCP. No API key.2368 PyPI2MIT