Skip to main content
Glama
Treeweft

treeweft-mcp

Official
by Treeweft

Treeweft is a GraphRAG code search engine that understands code structurally, not as flat text. It parses every file into an abstract syntax tree via tree-sitter, then weaves those parsed pieces into a connected graph in Neo4j, where calls, imports, inheritance, and references become threads connecting the nodes.

Search traverses that woven structure — semantically through vector embeddings and structurally through the graph — instead of grepping flat text. In a September 2026 benchmark on a pinned production codebase, across three current agent models, an agent using treeweft matched the answer quality of a grep/glob/read agent at 11–47% lower cost — with the largest savings on the questions where you don't know the name of what you're looking for. See Benchmarks.

Treeweft was previously named Treeloom. Upgrading an existing install? See Migrating from Treeloom. Compatibility with Treeloom names ends on 2026-11-01.

How It Works

Source Code  ──►  tree-sitter  ──►  Abstract Syntax Trees
                                          │
                                   (parse the trees)
                                          │
                                          ▼
                                    Neo4j Graph DB
                                          │
                                   (weave the graph)
                                          │
                                          ▼
                              Milvus Vector Store  ──►  Search API
  1. Parse: tree-sitter parses every file into AST chunks — functions, classes, blocks, with their hierarchical relationships preserved

  2. Embed: Text chunks are embedded via TEI (Jina, BGE, etc.) into Milvus vector store

  3. Weave: Neo4j captures calls, imports, inheritance, and references as graph edges between parsed nodes

  4. Search: Hybrid retrieval combines dense vector search, BM25 keyword matching (Milvus and LanceDB), graph-enhanced scoring, and HyDE query expansion

What is HyDE?

Hypothetical Document Embeddings (Gao et al., 2022). Instead of embedding a raw query like "how does auth middleware work?" — which may share little vocabulary with the actual code — Treeweft asks the LLM to hallucinate a short code snippet that would answer the question, then embeds that. The hypothetical code's embedding is far closer to real matching code in vector space than the raw query would be. It's query expansion via imagination.

Related MCP server: semhood

Quick Start

Simple mode (no GPU, one API key)

Evaluating on a laptop? The simple profile swaps Milvus → embedded LanceDB, Neo4j → embedded SQLite, and TEI/local-LLM → the OpenAI API (embeddings + HyDE/summaries), with reranking on the CPU. One Docker container (Postgres) by default.

git clone https://github.com/Treeweft/treeweft.git && cd treeweft
cp .env.simple.example .env   # then paste your OPENAI_API_KEY into it
./run-simple.sh

See docs/simple-mode.md for the full walkthrough, cost expectations, and the capability matrix vs the full stack.

Full stack

Prerequisites: Docker (with compose), Python 3.11+, and an OpenAI-compatible LLM endpoint for HyDE/summaries — the .env.example default expects Ollama on the host (ollama pull qwen2.5-coder:7b); if the LLM is unreachable, indexing and search still work, just without HyDE and chunk summaries.

git clone https://github.com/Treeweft/treeweft.git
cd treeweft
cp .env.example .env   # quickstart defaults: everything local

# One-shot: pip install -e ., build mcp-server (image only — not started by
# default), bring up Postgres + TEI + Neo4j + Milvus
# (COMPOSE_PROFILES=local-infra), and run the host-side indexer in the
# foreground on :8001. MCP access is stdio (treeweft-mcp) by default, no
# container; add http-mcp to COMPOSE_PROFILES for the deprecated HTTP+SSE
# transport.
./run.sh

No NVIDIA GPU? Add cpu to COMPOSE_PROFILES in .env and set EMBEDDING_URL=http://localhost:8083 — embedding runs on the CPU TEI image (slower indexing, identical results). Already running Neo4j/Milvus elsewhere? Drop local-infra from COMPOSE_PROFILES and point NEO4J_URI / MILVUS_URI at your instances.

If you'd rather start the pieces by hand:

pip install -e .

# Docker side: Postgres, TEI embedding + reranker, Neo4j, Milvus
docker compose up -d postgres tei-embedding tei-reranker neo4j milvus

# Host side: the indexer must NOT run in Docker — it reads arbitrary
# user paths off the host filesystem.
uvicorn treeweft.indexer_service:app --host 127.0.0.1 --port 8001

The MCP server is not a compose service by default: it runs as a stdio subprocess spawned by your MCP client (treeweft-mcp). The deprecated HTTP+SSE transport is available with COMPOSE_PROFILES=http-mcp.

Prebuilt images for the Treeweft services (treeweft/indexer, treeweft/mcp-server, treeweft/ui, treeweft/qwen3-reranker) are published to Docker Hub on every release; docker compose pull fetches them instead of building. See docs/docker-images.md for tags, platforms, and the release procedure.

Point your MCP client at the treeweft-mcp console script — it reaches the host indexer directly via INDEXER_URL (no host.docker.internal needed on stdio):

{
  "mcpServers": {
    "treeweft": {
      "command": "treeweft-mcp",
      "env": {
        "INDEXER_URL": "http://localhost:8001",
        "TREEWEFT_MCP_TOKEN": "<your treeweft PAT or API key>"
      }
    }
  }
}

Once everything is up:

# Index a repo (host indexer, :8001)
curl -X POST http://localhost:8001/index-directory \
  -H 'Content-Type: application/json' \
  -d '{"directory": "/path/to/your/repo", "pattern": "**/*"}'

# Search (host indexer, :8001). A scope is required: path_prefix for one
# repo, or "cross_repo": true to deliberately search the whole index.
curl -X POST http://localhost:8001/search \
  -H 'Content-Type: application/json' \
  -d '{"query": "how does auth middleware work?", "path_prefix": "/path/to/your/repo"}'

Before kicking off a long index, run preflight to predict wall time and get tuning hints:

python -m treeweft.preflight /path/to/repo
# or hit the endpoint. A `path` scan is refused (403) unless it sits under
# PREFLIGHT_ALLOWED_ROOTS (colon-separated) in the indexer's environment:
curl -X POST http://localhost:8001/preflight \
  -H 'Content-Type: application/json' \
  -d '{"path": "/path/to/repo"}' | python -m json.tool

Operator UI: the ui/ standalone SPA (see ui/README.md) gives you Jobs, Sources, and Submit tabs against the indexer. Run it with npm install && npm run dev from ui/. When AUTH_ENABLED=true, log in with a bearer token (personal access token or API key); it's stored in localStorage.

When AUTH_ENABLED=true, the indexing endpoints (/index-*), GET /sources, and search (/search, /graph-explore) require -H "Authorization: Bearer $TREEWEFT_KEY"; search is also checked against per-repo grants. /health, /status, and GET /jobs stay open (job views are redacted for anonymous callers). Set TREEWEFT_SEARCH_OPEN=1 to keep /search anonymous while indexing stays protected (/graph-explore still needs a token).

Architecture

┌─────────────────────────────────────────────┐
│                  MCP Server                  │
│             (stdio, per client)              │
│          Pure API consumer — zero DB access   │
└──────────────┬──────────────────────────────┘
               │ HTTP
┌──────────────▼──────────────────────────────┐
│               Indexer Service                │
│              (Host :8001)                    │
│    Owns all data access: Milvus, Neo4j, TEI, │
│    PostgreSQL, tree-sitter, embedding        │
└──────┬──────────────┬───────────┬───────────┘
       │              │           │
  ┌────▼────┐   ┌─────▼─────┐  ┌─▼──────────┐
  │ Milvus  │   │   Neo4j   │  │ PostgreSQL  │
  │(vector) │   │  (graph)  │  │ (persistence)│
  └─────────┘   └───────────┘  └─────────────┘

DDD layers: domain/ (pure logic) → adapters/ (I/O implementations) → application/ (FastAPI wiring) → infrastructure/ (config, DI). Domain never imports adapters.

Features

  • Structural parsing: tree-sitter AST chunking preserves function/class/block hierarchy

  • Graph weaving: Neo4j captures calls, imports, inheritance, references

  • Hybrid search: Dense vectors + BM25 + graph-enhanced scoring + HyDE
    (BM25 available with Milvus and LanceDB; ChromaDB uses embedding-only hybrid search)

  • Dual GPU/CPU embedding: Token-cost-aware routing keeps oversized chunks off the GPU, falls back to CPU

  • Multi-vector backend: Milvus (default), LanceDB, ChromaDB — selected with VECTOR_STORE
    (non-default backends are install extras: pip install -e .[lancedb], .[chromadb], or .[all-backends])

  • Community detection: Louvain algorithm identifies code clusters

  • Webhook indexing: Auto-reindex on merged pull requests (GitHub and Gitea webhooks)

  • Auth & authorization: API key auth, per-repo access grants, groups

  • MCP protocol: Standard Model Context Protocol server for AI agent integration

Benchmarks

Short answer. Across three current agent models, an agent using treeweft matched the answer quality of a grep/glob/read agent at 11–47% lower cost, with fewer turns, in every configuration we measured. The cost win is statistically significant per query on symbol-free questions for every model. Treeweft's answer-quality score was higher in 5 of 6 configurations, but not significantly so in any — so the claim is the same answers for less, not better answers.

The benchmark is agentic: a real LLM agent answers 100 curated questions about featbit — a ~1,900-file C#/TypeScript production codebase, pinned at tag v5.4.9 — once per retrieval approach, with identical prompts and a 12-turn cap. Answers are scored against gold references by an independent judge (gpt-4o-2024-08-06, not any of the agents). Two question populations run separately: symbol-bearing (the identifier is in the question) and symbol-free (a behavioural description with no name to search for).

Treeweft vs the grep/glob/read (Claude-Code-style) agent, September 2026, n=100 per row:

agent model

questions

cost vs grep (per-query p)

turns (grep → treeweft)

12-turn caps (grep / treeweft)

correctness (grep → treeweft, p)

deepseek-flash

symbol-bearing

−11% (p=0.057)

4.7 → 3.6

0 / 0

4.52 → 4.49 (p=0.851)

deepseek-flash

symbol-free

−30% (p=5.6e-07)

6.7 → 4.2

11 / 4

4.40 → 4.56 (p=0.359)

claude-sonnet-5

symbol-bearing

−14% (p=0.00041)

4.8 → 3.3

2 / 1

4.41 → 4.47 (p=0.286)

claude-sonnet-5

symbol-free

−25% (p=4.7e-06)

6.2 → 3.3

11 / 3

4.14 → 4.43 (p=0.185)

gpt-5.5

symbol-bearing

−20% (p=0.035)

5.5 → 5.0

5 / 6

4.30 → 4.38 (p=0.345)

gpt-5.5

symbol-free

−47% (p=3.1e-17)

8.2 → 5.2

24 / 8

4.05 → 4.41 (p=0.061)

What the numbers say:

  • Treeweft is cheaper. The per-query cost difference is significant in 5 of 6 rows (deepseek-flash on symbol-bearing questions is borderline, p=0.057) and on every symbol-free row. The saving is largest on the most expensive model, where every extra exploratory turn costs more.

  • On symbol-free questions grep runs out of road. It hit the 12-turn cap on 11–24% of questions, significantly more often than treeweft on every model. Capable agents usually still reach the answer; they just pay for the search.

  • On symbol-bearing questions grep is hard to beat at finding the file — it string-matches the identifier. Ranking is a statistical tie on two models, and grep ranks the right file first significantly more often on deepseek-flash (p=0.023). Treeweft still gets there for less.

  • vs vector RAG with identical embeddings (zilliztech/claude-context, deepseek-flash): treeweft ranks the right file first significantly more often — recall@1 0.68 vs 0.48 (p=0.0017) on symbol-bearing and 0.44 vs 0.28 (p=0.007) on symbol-free questions — at 11–22% fewer tokens, with answer quality tied.

Full tables, per-query significance, setup, spend, and reproduction steps: docs/benchmark-results-2026-09.md.

Semantic search, not symbol lookup

What you ask matters more than which tool you use. When the identifier is already in the question, grep finds it directly and treeweft's advantage is cost alone (11–20% cheaper). When you can only describe the behaviour, grep has nothing to match and flounders; treeweft's cost advantage widens to 25–47%. Don't use a semantic engine for symbol lookup — grep already nails it. Use treeweft for "how does this work?" and "where is the code that does X?".

The graph layer earns its keep mainly through the neighbor and community context it feeds the agent. Whether graph rescoring of results adds anything on top depends on the strength of the reranker — see docs/benchmark-findings.md and docs/graph-ablation.md.

Earlier results (June 2026)

An earlier multi-repo run — featbit plus PowerToys (C#), guava (Java) and three.js (JS), n=100 per repo, agent deepseek-chat — found treeweft significantly better than grep on symbol-free answer quality (pooled p<0.0001), with the grep agent collapsing outright on PowerToys (recall 0.27, 43k tokens of failed searches). That run also measured aider's repo-map (an 8k-token symbol map re-sent every turn: 5.1× grep's tokens, and worse retrieval — see docs/benchmark-refresh-runbook.md §1) and a cost comparison including Claude Opus 4.8 (−16% vs grep). Two things limit how far to lean on it: it used an earlier version of the benchmark harness whose numbers this project does not mix with current ones, and its significant answer-quality edge did not replicate on featbit with current models, whose agents keep grep competitive. The cost direction did replicate.

That pattern — a quality edge with weaker agents, cost-only with stronger ones — shows up within the current harness too: a July 2026 run with gpt-4o as the agent (weaker than the models above, and judging its own answers) found treeweft significantly better on symbol-free answer quality on featbit (p=2.3e-10), with grep capping out on 26% of questions. When grep has nothing to match, a weaker agent gives up; a capable one keeps searching and pays for it. Details: docs/benchmark-findings.md and docs/cost-across-models.md.

Caveats, stated plainly

  • One repository in the September run (featbit: C# back end, TypeScript front end), and one run per configuration — no repeated-run variance estimate.

  • The judge scale is compressed. Correctness means sit between 4.05 and 4.56 out of 5, with 71–81 ties per 100 questions, which makes small quality differences hard to detect.

  • Recall counts files the agent opened, so an arm that takes more turns scores higher on it mechanically; read grep's recall numbers with that in mind.

  • A treeweft search takes roughly 3 seconds, most of it reranking — latency traded for fewer agent turns.

  • Raw per-query result rows are not shipped in the repo; the query sets and gold answers are. Reproduce with python -m treeweft.benchmark agentic --help.

We also ran a milestone of token-optimization experiments trying to shrink the search payload, and rejected nearly all of them: a smaller response just makes the agent re-read (compensatory fetch), raising turns and erasing the saving. The only default-on win was serialization (markdown encoding, −26%), not content. The methodology and every rejected experiment — including across five agent models, where trimming gets worse as capability rises — are written up in docs/token-optimization-rejected.md.

Configuration

Copy .env.example to .env and configure:

Variable

Default

Purpose

EMBEDDING_MODEL

jinaai/jina-embeddings-v2-base-code

TEI embedding model

RERANKER_MODEL

BAAI/bge-reranker-v2-m3

TEI reranker model

VECTOR_DIM

768

Must match embedding model

LLM_URL

http://localhost:11434/v1

OpenAI-compatible LLM for HyDE/summaries (Ollama default)

DATABASE_URL

required

Postgres connection string (job state + summary cache)

See .env.example for all options.

Observability

Treeweft emits OpenTelemetry traces from the indexer for every job and per-stage span (parse_and_chunk, embed.chunks, summaries.*, milvus.*, graph.index_file). The bundled docker-compose.yml defines an OTel collector (head sampling, always_on by default — tail sampling belongs in a central collector, see docs/observability-runbook.md), Tempo and Grafana; run.sh does not start them, so bring them up with docker compose up -d otel-collector tempo grafana. Point Grafana at the Tempo datasource (assets/observability/grafana-datasources.yaml) and import assets/grafana/treeweft-indexer-traces.json for the pre-built per-stage dashboard. Configure via the standard OTEL_EXPORTER_OTLP_ENDPOINT / OTEL_TRACES_SAMPLER env vars; span names use the stable treeweft. attribute prefix. Set OTEL_SDK_DISABLED=true to turn it off entirely.

Testing

# The bare `pip install -e .` from Quick Start above is not enough — the
# full suite also exercises every optional-backend adapter. Pull in dev
# (pytest) + all-backends + simple:
pip install -e ".[dev,all-backends,simple]"
pytest tests/unit -q     # 1,656 Detroit-style unit tests (no services needed)

Retrieval quality is validated separately against a live stack via the benchmark harness (python -m treeweft.benchmark run|ab|agentic).

Documentation

docs/index.md maps every document under docs/: a one-line summary of each, how far to trust it, and a recommended reading order for evaluating, operating, and changing Treeweft.

Brand

See assets/branding/ for logo files, favicon, and full brand guidelines. The Thread-Tree mark is the canonical Treeweft symbol.

License

MIT

Available Tools

20 tools
explain_codeA

Use when the user asks you to explain a specific concept or section of code and wants tight, focused context (e.g. 'explain how X works' rather than 'find everything about X'). Same as search_code but with top_k=5 and an optional file_path filter applied server-side — useful for narrowing to a single file you already know is relevant.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
languageNo
use_hydeNo
file_pathNo
source_idNo
cross_repoNo
use_hybridNo
path_prefixNo
rerank_poolNo
use_graph_scoringNo
use_summary_vectorNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It discloses key behavioral traits: top_k=5, optional file_path filter applied server-side, and that it returns focused context. However, it does not describe the output format or any side effects, which is a minor gap but not misleading.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with usage context and key differentiators. Every word earns its place, and it avoids unnecessary detail while conveying the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (11 parameters, no output schema, no annotations), the description is incomplete. It does not explain what the tool returns, how to use most parameters, or any limitations beyond the top_k and file_path filters. An agent would struggle to use this tool effectively with only this description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It only mentions top_k (implicitly) and file_path, leaving 9 other parameters (language, use_hyde, cross_repo, etc.) unexplained. This is insufficient for a tool with 11 parameters, as the agent has no guidance on how to set them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear purpose: explaining a specific concept or section of code with tight, focused context. It explicitly contrasts with search_code, saying it is the same but with top_k=5 and an optional file_path filter, which distinguishes it from siblings and clarifies its role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use guidance: use when the user wants a focused explanation (e.g., 'explain how X works') rather than broad search. It names the alternative (search_code) and highlights the differences, giving the agent a clear routing decision.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_callersB

Find callers of a function/method. Accepts either a name (auto-resolves; returns callers across all matches with target_id set) or a full entity id.

ParametersJSON Schema
NameRequiredDescriptionDefault
source_idNo
name_or_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explains that name resolution auto-resolves and returns callers across all matches with target_id set, which is a useful behavioral detail. However, it does not disclose whether the operation is read-only, any potential side effects, error behavior, or what happens when no callers are found or the entity does not exist. This is insufficient for a query tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense sentence with no filler. It front-loads the core purpose ('Find callers of a function/method') and packs the two input modes into a compact, readable structure. Every phrase earns its place, making it both concise and efficiently structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists and the parameter count is small, the description leaves critical gaps: source_id is unexplained, and there is no guidance on when to use this tool relative to siblings like find_references. The tool could be called correctly with name_or_id alone, but an agent lacks context about optional parameters and selection criteria, making the description incomplete for optimal usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the missing parameter documentation. It explains the meaning of name_or_id (either a name that auto-resolves or a full entity id) and notes that results include target_id. However, source_id is completely undocumented in both the schema and the description, leaving agents to guess its purpose and behavior. The description adds some value for one parameter but fails to cover the other.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: finding callers of a function/method. It specifies the resource (function/method) and the action (find callers), and distinguishes it from sibling tools like find_references by focusing on callers specifically. The two input modes (name or full entity id) are explicitly described, making the intent unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by defining what the tool does and how inputs are handled, but it provides no explicit guidance on when to use this over alternatives like find_references or graph_explore. There is no mention of scenarios where this tool is preferable or when it should not be used. The input mode details (auto-resolve name vs. full id) are helpful but not contextual usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_definitionA

Find where a symbol is defined. Looks up Class/Function entities by name. Returns all matches across files (let the caller disambiguate). Optional kind filter ('Class' or 'Function') and source_id to scope to one indexed source.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
nameYes
source_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden and discloses key behavior: it returns every match across files and leaves disambiguation to the caller, indicating a read-oriented lookup. It doesn't explicitly state side-effect-freeness, but the 'Find/Looks up/Returns' framing makes the read-only intent clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences front-load purpose, then behavior, then filters; no filler. The only minor redundancy is 'Find where a symbol is defined' vs 'Looks up ... by name', but it doesn't hurt clarity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the 3-param lookup complexity and an available output schema, the description supplies enough invocation context: what is searched, what is returned, and what each parameter does. A statement about no side effects or source_id provenance would be a small addition, but is not essential.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so this dimension relies on the description. It defines name (lookup by name), kind (allowed values 'Class' or 'Function'), and source_id (scope to one indexed source), which covers all three parameters' primary semantics, though it doesn't specify name format or case sensitivity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names the exact operation ('Find where a symbol is defined') and the second narrows the resource to Class/Function entities, which clearly separates this from sibling tools like find_references and find_callers. It also states the return policy (all matches, caller disambiguates), so there's no ambiguity about scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear context: use when you need definition locations for a named Class/Function, and use optional filters to narrow by kind or source. It does not explicitly name alternatives or when-not conditions, but its purpose is specific enough to route selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_referencesA

Find all references (any incoming relationship — CALLS, INHERITS, IMPORTS, DEFINES) to a symbol. Accepts a name or a full entity id.

ParametersJSON Schema
NameRequiredDescriptionDefault
source_idNo
name_or_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses that all incoming relationship kinds are returned and that input can be a name or entity id, but it does not mention behavior on ambiguous names, result limits, or whether the operation is read-only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences with the core behavior and relationship taxonomy front-loaded. There is no filler or redundant restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the basic call is feasible, but the description leaves source_id unexplained and gives no guidance for choosing among the many graph/reference sibling tools. It is adequate but not complete for confident selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies name_or_id as accepting 'a name or a full entity id', but says nothing about the optional source_id parameter, leaving its purpose unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Find all references' to a symbol, and enumerates the exact incoming relationship types (CALLS, INHERITS, IMPORTS, DEFINES). This clearly distinguishes it from siblings like find_definition and find_callers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for discovering all incoming references, but it does not explicitly say when to choose this over find_callers, find_definition, or search_code. No exclusions, prerequisites, or alternative routing guidance is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_index_jobB

Get the current state of an index job (queued/running/done/failed).

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. The word 'Get' implies a read-only operation, which is a positive signal. However, the description does not disclose any potential side effects, error conditions (e.g., invalid job_id), or whether the operation is safe. For a simple getter, the description provides minimal behavioral context beyond the name.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that front-loads the action ('Get') and the resource ('current state of an index job'). It is efficiently worded with no redundant information, making it easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with one parameter and no output schema. The description lists the possible states, which tells the agent what to expect in the response. It does not mention error handling or the exact return format, but for a basic getter this is sufficient. No critical missing context for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does not explicitly mention job_id or its role. The description only says 'an index job', which indirectly refers to the parameter but does not explain its format, constraints, or how it relates to the job. The schema only lists the parameter as a string with a title, so the description adds little semantic value for the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Get' and the resource 'current state of an index job', and lists the possible states (queued/running/done/failed). It distinguishes this tool from siblings like list_index_jobs (which lists all jobs) and wait_for_index_job (which blocks). However, it does not explicitly name an alternative, so it loses a point for not being more explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use this tool (when you have a job_id and want its current state), but it does not provide explicit guidance on when not to use it or mention alternatives. Given the sibling tools, an agent might confuse it with list_index_jobs or wait_for_index_job. No exclusions or alternative routing is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

graph_exploreA

Use when the user wants a structural/architectural view rather than a code-snippet view — 'what depends on X', 'what does X talk to', 'show me the call graph around Y'. Finds entities related to the query and traverses their connections (callers, callees, imports, inheritance) up to depth hops. For 'find the code that does X', use search_code instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
depthNo
queryYes
languageNo
use_hydeNo
source_idNo
use_hybridNo
path_prefixNo
rerank_poolNo
use_graph_scoringNo
use_summary_vectorNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the core behavior: finds entities related to the query, then traverses callers, callees, imports, and inheritance up to `depth` hops. This goes beyond the title and gives the agent a clear model of what happens. It stops short of stating whether the operation is strictly read-only or how results are returned, but the exploratory nature is well conveyed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two tight sentences. The first front-loads the use case with concrete examples, and the second states the traversal behavior while naming the alternative tool. Every clause adds value, with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 10 parameters, no description coverage in the schema, no annotations, and no output schema. The description covers when to use it and the core traversal idea, but it omits the meaning of most parameters, return format, and any caveats or prerequisites. For a tool of this complexity, this is a material gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all 10 parameters, but it only explains `depth` ('up to depth hops'). Parameters like `use_hyde`, `use_hybrid`, `rerank_pool`, `use_graph_scoring`, and `use_summary_vector` are left entirely to name inference, which is insufficient for a tool with this many tuning knobs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('explore', 'traverses') and resource (graph/connections) and clearly states the tool provides a structural/architectural view rather than code snippets. It explicitly distinguishes itself from `search_code`, and the examples ('what depends on X', 'call graph around Y') make the purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance: when the user wants dependency/call-graph/architecture insight. It also provides an explicit exclusion and alternative: for 'find the code that does X', use `search_code` instead. This leaves little ambiguity about tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

hydrate_chunksA

Fetch the full code bodies for facet search hits. After a search_code call with response_mode='facet' (which returns ranked metadata + headers but NO code), pass the hit_id values of the results you care about here — in ONE batched call — to get their exact code bodies. Cheaper and more precise than read_file: it returns exactly the indexed chunk(s), reconstructing merged ranges. Returns compact markdown by default; pass response_format='json' for the structured dict.

ParametersJSON Schema
NameRequiredDescriptionDefault
idsYes
source_idNo
response_formatNomarkdown

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses that facet search returns no code, that this tool returns exact indexed chunks, reconstructs merged ranges, and supports markdown or JSON output. It does not cover error behaviors or edge cases, but for a read-oriented hydration tool it provides substantial behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded: the first sentence states the core action, the second gives the workflow context, and the third covers output options. Every sentence contributes useful information without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for the core calling pattern: it identifies the prerequisite call, the required input, the batching behavior, and the output format options. The main gap is the unexplained `source_id` parameter and the lack of guidance on failure or stale hit_id scenarios, but the output schema helps cover return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains that `ids` should be facet hit_ids and that `response_format` accepts 'markdown' (default) or 'json'. However, the `source_id` parameter is never mentioned, leaving one optional parameter semantically unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: fetch full code bodies for facet search hits. It differentiates itself from alternatives by referencing the specific search_code facet flow and read_file comparison. The verb 'fetch' and resource 'code bodies' are concrete and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells the agent when to use this tool: after a search_code call with response_mode='facet', passing the hit_id values. It also names an alternative (read_file) and explains why hydrate_chunks is cheaper and more precise, giving clear decision criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

index_directoryA

Index all code files in a directory (recursive). Returns immediately with a job_id; poll get_index_job to track progress. Set force=true to re-index even if the directory is already indexed and unchanged (otherwise the call no-ops to the prior job).

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNo
patternNo**/*
directoryYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does well: it discloses async return, the job_id, and the force re-index/no-op behavior. It does not mention side effects (index persistence) or permissions, but the core behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences, front-loaded with the primary purpose. Every sentence adds value: purpose, async behavior, and force semantics. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an async tool, it covers the key workflow: immediate return, job_id, and polling via get_index_job. It does not explain the pattern parameter or error handling, but given the simplicity and no output schema, it is reasonably complete. The missing pattern is the main gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain parameters. It explains force ('Set force=true to re-index... otherwise no-ops'), and directory is implied by the title. However, the pattern parameter is never mentioned, leaving its meaning and default unexplained, which is a significant gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action and resource: 'Index all code files in a directory (recursive)'. This distinguishes it from siblings like index_file (single file) and index_repo (repo-level). The recursive scope and directory target are explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context: it is asynchronous (returns immediately with job_id, poll get_index_job), and explains the force parameter's no-op behavior. However, it does not explicitly exclude alternatives such as index_file for single files or mention when to prefer another tool, leaving some inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

index_fileA

Index a single code file. Parses with tree-sitter, chunks with chonkie, embeds via TEI, and stores vectors in Milvus + entities in Neo4j. Returns immediately with a job_id; poll get_index_job to track progress. Set force=true to re-index even if this file is already indexed and unchanged (otherwise the call no-ops to the prior job).

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNo
file_pathYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does so thoroughly. It discloses that the tool is asynchronous ('Returns immediately with a job_id'), names the follow-up tool for tracking, explains the no-op behavior for unchanged files, and details the underlying pipeline. This is rich behavioral context beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, zero filler, and the core action is front-loaded. Each sentence adds necessary information: what it does, how to track it, and when to adjust behavior. The structure is ideal for an AI agent scanning quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two simple parameters, no annotations, and no output schema, the description is complete. It covers the asynchronous workflow, the no-op case, the force option, and the downstream storage destinations. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains force's meaning and effect in detail ('re-index even if this file is already indexed and unchanged'). file_path is not elaborated beyond 'single code file,' but its name and string type make it self-explanatory. This is good compensation for a small parameter set.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: 'Index a single code file,' which clearly differentiates it from siblings like index_directory and index_repo. It also states the full processing pipeline and the asynchronous result behavior, leaving no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: this is for a single file, implying index_directory/index_repo are alternatives for broader scopes. It also explains when to set force=true and when the call no-ops, which is concrete usage guidance. It does not explicitly enumerate sibling alternatives or exclusion criteria, but the single-file scope is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

index_graphA

Build (or rebuild) the Neo4j code graph for an already-indexed source — pass 2 of two-pass indexing. Use this when chunks/vectors exist but graph-backed tools (find_definition, find_callers, find_references, graph_explore) return empty, or after a chunks-only index (skip_graph=true). Drops the source's existing graph entities and re-extracts; safe to re-run. Returns immediately with a job_id; poll get_index_job. Get source_id from list_indexed_sources.

ParametersJSON Schema
NameRequiredDescriptionDefault
source_idYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly warns that it 'Drops the source's existing graph entities and re-extracts' – a destructive side effect – and further states it 'Returns immediately with a job_id; poll get_index_job', revealing the asynchronous nature. It also notes 'safe to re-run', which indicates idempotency. This is a thorough behavioral disclosure for a tool with no annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the first line states the purpose, then usage conditions, then behavior, then async handling, and ends with a concrete sourcing instruction. Every sentence adds value; no filler or repetition. It packs a lot of essential context into two sentences without becoming unwieldy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with one parameter, no annotations, and no output schema, the description covers all necessary aspects: what it does, when to use it, potential destructive behavior, async nature, how to track progress, and where to obtain the source_id. There is no missing information an agent would need to invoke or monitor this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema only declares source_id as a required string with no description coverage, leaving the description to compensate. The description adds crucial meaning by indicating that source_id refers to an already-indexed source and explicitly instructs 'Get source_id from list_indexed_sources'. This goes beyond the schema, though it doesn't explain the ID format or any validation rules, but for a single param this is adequate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Build (or rebuild)') and resource ('Neo4j code graph for an already-indexed source'), and immediately identifies this as 'pass 2 of two-pass indexing'. It distinguishes itself from siblings by explicitly referencing the graph-backed tools that return empty and by noting it operates on an already-indexed source, which sets it apart from index_file/index_directory/index_repo.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: 'when chunks/vectors exist but graph-backed tools return empty, or after a chunks-only index (skip_graph=true)'. Provides a prerequisite for obtaining source_id from list_indexed_sources and mentions the alternative of polling via get_index_job. This gives clear decision criteria and action steps.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

index_repoA

Index a git repository. Accepts a local path (with ~/ expansion) or a remote git URL. Auto-detects current branch for local repos. Returns immediately with a job_id; poll get_index_job to track progress. Re-indexing the same source upserts (no duplicates). By default an already-indexed, unchanged source no-ops to its prior job — set force=true to force a full re-index.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNo
pathNo
forceNo
branchNo

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure and does so thoroughly: it reveals the async return via job_id, the polling mechanism, upsert/no-duplicate semantics, and the no-op behavior for unchanged sources unless force=true. This goes well beyond the basic 'index' action.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each adding distinct value: what the tool does, how inputs are provided, and what behavioral results to expect. The core purpose is front-loaded and every sentence earns its place with no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the job_id lifecycle, re-indexing behavior, force re-indexing, and acceptable source types. It does not explicitly state whether url and path are mutually exclusive or mention auth requirements, but for a 4-parameter tool with no output schema, it is largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains path vs URL usage, the force flag's effect, and branch auto-detection for local repos. The branch parameter's role is only implied as an override rather than explicitly stated, which is a minor gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence 'Index a git repository' gives a specific verb and resource. The rest of the description further distinguishes it from siblings like index_file and index_directory by clarifying that it handles git repositories via local path or remote URL.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context on input types (local path with ~/ expansion or remote git URL) and the asynchronous workflow (returns job_id, poll get_index_job). It does not explicitly name alternative indexing tools or state when not to use this tool, but the usage context is strong enough to guide an agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_indexed_sourcesA

List all indexed sources (repos, directories, files) with their stats.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only states that it lists sources and their stats, but does not reveal any potential side effects, performance implications, or whether the listing is read-only. For a no-parameter tool, the description is minimal and leaves the agent without information about ordering, pagination, or potential cost.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence that front-loads the action ('List all indexed sources') and specifies the scope and content. There is no extraneous information, and every word serves a purpose. It is optimally concise for a simple listing operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so the description need not explain the return format. The description states the tool lists all indexed sources with their stats, which is sufficient for a no-parameter tool. Minor details like whether the list is paginated or ordered are not addressed, but the output schema likely covers the structure. Overall, it is reasonably complete for its simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema coverage is trivially 100%. According to the baseline for tools with no parameters, a score of 4 is appropriate. The description does not need to explain parameters that do not exist, and no additional semantic clarification is required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists all indexed sources (repos, directories, files) with their stats. It uses a specific verb and resource, and it is distinct from sibling tools like list_index_jobs or the indexing operations. The agent can unambiguously identify this as the listing tool for indexed sources.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention any exclusions, prerequisites, or typical use cases. There is no direction like 'use this to view all indexed sources' or 'instead of X'. The agent must infer its purpose solely from the name and description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_index_jobsA

List indexing jobs across all sources. Each entry includes status (queued/running/done/failed), kind (file/directory/repo/graph/incremental), processed_files / total_files, errors, current_file, and timing. Useful for answering 'is treeweft busy?' or 'what happened to my last index job?'.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the behavioral disclosure burden. It does disclose scope ('across all sources') and what each entry contains, which is useful. However, it does not address filtering behavior, ordering, limits, or whether this is a purely read-only snapshot; those gaps keep it at minimum viable rather than strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two focused sentences: the first states the core action and return contents, the second gives concrete usage questions. No filler or redundant schema repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is an output schema, and the tool is simple with only one optional parameter. Still, the complete absence of parameter documentation is a real gap, and the description would be stronger if it mentioned that status filters results and what values are accepted. Overall adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has one optional 'status' parameter with 0% description coverage, and the description never mentions that status can be used as a filter. It only lists status as an output field. Since schema coverage is low and the description does not compensate, an agent gets little guidance on how or why to supply this parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource ('List indexing jobs') and clearly defines scope ('across all sources'). It enumerates the fields returned and includes concrete user questions, making the tool's purpose unmistakable even among many sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear use cases ('is treeweft busy?', 'what happened to my last index job?'), which tells an agent when to select this tool. It does not explicitly contrast it with get_index_job, wait_for_index_job, or list_job_errors, so exclusions/alternatives are not fully spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_job_errorsA

List the per-file errors recorded during a specific indexing job. Each entry includes file_path, error_kind ('timeout' or 'exception'), error_message, elapsed_s, and occurred_at. Useful for 'which files failed?' and 'why?' without grepping process logs.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
job_idYes
offsetNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the return structure in detail (file_path, error_kind, error_message, elapsed_s, occurred_at) and the read-only nature via the verb 'List.' It does not mention pagination behavior or how missing job IDs are handled, but the core behavioral traits are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The first sentence front-loads the main action and enumerates the output fields; the second adds a practical use case. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple read-only nature, the description covers the purpose, the returned entry fields, and the intended use case. Pagination defaults and required parameters are already present in the schema, and the absence of an output schema is compensated by explicitly listing the return fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies job_id as the identifier of a specific indexing job, but says nothing about limit/offset semantics or their defaults. The description adds useful meaning for the primary parameter but not for the pagination parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'List the per-file errors recorded during a specific indexing job.' The resource is clearly distinct from siblings like list_index_jobs or get_index_job, and the 'which files failed?' and 'why?' clause reinforces the concrete use case.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear use case: 'Useful for which files failed? and why? without grepping process logs.' This implies when to use the tool but does not explicitly name alternatives or exclusion conditions, so it falls short of the full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rebuild_all_graphsA

Bulk-build the Neo4j code graph for indexed sources. Lists sources, then enqueues a graph job (POST /index-graph) for each. Use this once to repair a bulk import where chunks were indexed but the graph was never built (so find_definition/graph tools are globally empty). Graph extraction is idempotent (drops + rebuilds), so rebuilding a healthy source is harmless. Defaults to rebuilding ALL sources — the graph_indexed flag is NOT a reliable signal for the empty-graph bug (jobs marked themselves indexed while writing nothing), so only_missing=true (which filters to graph_indexed=false) will skip the very sources that need repair. Returns one entry per source with its job_id (or error). Jobs run through the shared queue — this kicks them off and returns; poll list_index_jobs(status='running') to track.

ParametersJSON Schema
NameRequiredDescriptionDefault
only_missingNo

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers. It discloses idempotence ('drops + rebuilds, so rebuilding a healthy source is harmless'), async queue behavior ('kicks them off and returns'), and the unreliable graph_indexed flag ('jobs marked themselves indexed while writing nothing'). It also specifies the return shape ('one entry per source with its job_id (or error)'), which is critical for a bulk async operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long (7 sentences) but each sentence earns its place: purpose, mechanics, trigger condition, idempotence, default with flag warning, return shape, and async follow-up. It could be trimmed slightly (e.g., 'POST /index-graph' is an implementation detail), but the density of relevant information remains high and the key caveats are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 1 optional parameter, no output schema, and no annotations, this description covers the decision to call it, the dangerous flag, the async nature, the return format, and the tracking mechanism. An agent can invoke it correctly and interpret results without needing additional documentation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate for the single boolean parameter. It does exactly that: 'only_missing=true (which filters to graph_indexed=false) will skip the very sources that need repair.' This adds critical semantic meaning beyond the Boolean type and default, explaining why the default is false and why setting it true is dangerous in the intended use case.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Bulk-build the Neo4j code graph for indexed sources.' It then clarifies the mechanism ('Lists sources, then enqueues a graph job for each') and ties it to a concrete failure mode ('find_definition/graph tools are globally empty'), which distinguishes it from single-source siblings like index_graph. The only_missing caveat further disambiguates correct vs. incorrect usage.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states the trigger condition: 'Use this once to repair a bulk import where chunks were indexed but the graph was never built.' It also warns against the harmful alternative ('only_missing=true ... will skip the very sources that need repair') and provides a follow-up action ('poll list_index_jobs(status='running') to track'), giving an agent a complete plan. No alternative tool is named, but the condition is precise enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_indexed_sourceA

Remove an indexed source. Drops its Milvus chunks and Neo4j entities (but preserves entities still referenced by other sources).

ParametersJSON Schema
NameRequiredDescriptionDefault
source_idYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It explicitly discloses the destructive side effects (drops Milvus chunks and Neo4j entities) and a safety nuance (preserves entities still referenced elsewhere). It does not explicitly state that the operation is irreversible, but the 'drops' language strongly implies permanence. This is adequate for a simple removal tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that front-loads the primary action and includes the key side-effect and preservation note. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that the tool has only one parameter and no output schema, the description covers the essential behavioral context: what gets removed and what is preserved. It does not mention potential errors or asynchronous behavior, but for a straightforward removal operation the description is reasonably complete. A slightly more complete version would note irreversibility, but the current description is sufficient for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides only the name 'source_id' with no description, and the tool description does not elaborate on what constitutes a valid source_id or how to obtain it. With 0% schema description coverage, the description should compensate, but it does not. The parameter semantics are left entirely to the parameter name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action (remove) on a specific resource (indexed source) and describes the primary effect (drops Milvus chunks and Neo4j entities). This clearly distinguishes it from sibling tools which are all about indexing, searching, or explaining code. The removal tool is unique among the siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when an indexed source needs to be removed, but it does not explicitly state when to use it versus other tools. There are no alternative removal tools, so no differentiation is needed, but it also does not mention any preconditions (e.g., source must exist) or post-conditions. It does mention a preservation nuance which gives some usage context, but overall it leaves the when-not-to-use implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_codeA

PRIMARY code-search tool — use FIRST for any question about how indexed code works, to plan a change, trace a flow, or find an implementation by behavior. Prefer this over grep/Read for indexed repos: grep needs exact strings, this finds code by meaning. Embeds the query, vector-searches Milvus, expands via Neo4j graph neighbors (callers/callees/imports/inheritance) and community summaries, reranks, returns chunks with file_path, start_line, end_line, snippet, score. Each chunk carries a citation field (e.g. 'repo@sha12:path') and the response includes a top-level sources map with commit_sha and permalink_base per source for building file-level permalinks. top_k defaults to 5 (the benchmarked sweet spot) — raise it only when a first search shows the answer spans many files. Optional filter: language (e.g. 'python'). SCOPE IS REQUIRED: pass source_id (from list_indexed_sources) or path_prefix (e.g. '/repo/src/') to scope to one repo, OR cross_repo=true to search the whole multi-repo index — exactly one, not both, or the search is rejected. Set check_staleness=true to add an is_stale flag per source (does live git HEAD resolution — slightly slower). response_mode (opt-in, default 'full'): 'facet' returns ranked metadata + a one-line header per hit with NO code body (cheaper — then fetch ranges with read_file); 'summary_tail' keeps the top-2 snippets and replaces lower-ranked bodies with their indexed one-line summary. Returns compact markdown by default (~20% fewer tokens); pass response_format='json' to get the structured dict instead (for programmatic callers).

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
top_kNo
languageNo
use_hydeNo
source_idNo
cross_repoNo
use_hybridNo
path_prefixNo
rerank_poolNo
adaptive_topkNo
response_modeNofull
strip_importsNo
check_stalenessNo
response_formatNomarkdown
adaptive_topk_gapNo
use_graph_scoringNo
use_summary_vectorNo
query_class_payloadNo
include_community_summariesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations to provide a safety profile, the description carries the full burden and does so thoroughly: it explains the embedding/vector-search/graph-expansion pipeline, what fields chunks contain, citation format, response modes, staleness checking with a performance cost, and output format. No annotation contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place; it is front-loaded with the most important usage information and then layers supporting detail. Dense structure with scoping rules and response modes, but no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Coverage is strong for common use cases: scope requirements, response modes, citations, and performance trade-offs are all present, and an output schema exists so return-value details are redundant. It loses one point because the many advanced schema parameters remain undocumented and there is no explicit distinction from the sibling search_code_enhanced tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds meaning to several core parameters (query, top_k, language, source_id, path_prefix, cross_repo, check_staleness, response_mode, response_format) beyond a schema that has 0% description coverage. However, 10+ parameters (use_hyde, use_hybrid, rerank_pool, adaptive_topk, strip_imports, include_community_summaries, etc.) are left completely unexplained, which is a clear gap for a tool with 19 parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a precise purpose: semantic code search over an indexed repository, returning ranked code chunks. It explicitly differentiates itself from grep/Read and establishes itself as the PRIMARY search tool, making the resource and behavior unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance ('use FIRST for any question about how indexed code works...') and names alternatives with the condition that selects them ('Prefer this over grep/Read for indexed repos: grep needs exact strings'). It also tells the agent when to raise top_k and warns against invalid scope combinations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_code_enhancedA

Use when search_code results suggest the answer spans multiple structurally-related files (e.g. handler/publisher/consumer triads, or a class and its subclasses) and you want to pull in chunks from those neighbor files in one shot. Same inputs and filters as search_code; does a second-pass Milvus search filtered to neighbor file paths before reranking. Results carry per-chunk citation fields and a top-level sources map (commit_sha + permalink_base). Set check_staleness=true to add an is_stale flag per source (does live git HEAD resolution — slightly slower). response_mode (opt-in, default 'full'): 'facet' returns ranked metadata + a one-line header per hit with NO code body (then fetch ranges with read_file); 'summary_tail' keeps the top-2 snippets and replaces lower-ranked bodies with their indexed summary. Returns compact markdown by default (~20% fewer tokens); pass response_format='json' to get the structured dict instead (for programmatic callers).

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
top_kNo
languageNo
use_hydeNo
source_idNo
cross_repoNo
use_hybridNo
path_prefixNo
rerank_poolNo
adaptive_topkNo
response_modeNofull
strip_importsNo
check_stalenessNo
response_formatNomarkdown
adaptive_topk_gapNo
use_graph_scoringNo
use_summary_vectorNo
query_class_payloadNo
include_community_summariesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it delivers: it discloses the two-pass search/rerank behavior, result citation fields and sources map, staleness checking cost, response_mode variants, and default output format. This goes well beyond a simple 'search' label.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but efficiently front-loaded: it opens with the decisive use condition, then packs behavioral and option details into a few sentences with no filler. Every sentence carries actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is complex (19 params, response modes, output formats) and the description covers most key behaviors and outputs, especially given an output schema exists. It leans on search_code for parameter definitions and doesn't spell out every parameter, but the essential invocation decisions (when, how, output shape) are addressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds real meaning for check_staleness, response_mode, and response_format, and references search_code for the rest. However, with 0% schema description coverage for 19 parameters, most inputs (e.g., top_k, use_hyde, cross_repo, rerank_pool) are left undefined in this tool's definition and only inherited by reference to a sibling.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific use case ('when search_code results suggest the answer spans multiple structurally-related files') and a specific action ('pull in chunks from those neighbor files in one shot'). It clearly distinguishes itself from the sibling search_code by describing the second-pass Milvus search and neighbor-file filtering.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly conditions use on search_code results and names the alternative directly. It also provides guidance among response modes, including when to fall back to read_file for 'facet' mode, and notes the performance tradeoff of check_staleness.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

source_stalenessA

Check whether an indexed source is out of date relative to its current upstream HEAD. For local-path sources, resolves via git rev-parse; for URL sources via git ls-remote. Returns {indexed_sha, current_sha, is_stale} — is_stale is null for non-git inputs.

ParametersJSON Schema
NameRequiredDescriptionDefault
source_idYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden. It discloses the underlying mechanism for local-path vs URL sources, the exact return shape, and the null behavior for non-git inputs. This exceeds the minimum expectation, though it stops short of mentioning error cases or network/auth dependencies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly crafted sentences: the action, the mechanism, and the return value. No filler, all information is relevant)Skip the previous line. The return type and edge case are front-loaded, making it efficient for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a single parameter fro no output schema and no annotations, the description is nearly complete. It defines the return object and covers the non-git edge case, but omits practical details like how to obtain a source_id or what happens when a local path or URL is unreachable. For such a simple tool, this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the single parameter. It reveals that source_id refers to 'an indexed source' and that it can be either a local-path or URL source, which adds meaning beyond the minimal schema. However, it does not explicitly address where source_id comes from or its expected format, leaving room for ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Check'), a specific resource ('indexed source'), and a precise outcome ('out of date relative to its current upstream HEAD'). It also distinguishes the tool from sibling indexing/search tools by describing a focused status-check behavior that none of the siblings cover.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool by explaining the check it performs)Skip the previous line. It does not explicitly name alternatives or state when not to use it. A user must infer that this is for verifying staleness rather than for indexing or searching, and there is no routing to sibling tools like list_indexed_sources.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

wait_for_index_jobA

Block until an index job reaches done/failed/cancelled, or until timeout_s elapses. Polls server-side. Use this when you want a synchronous index call.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes
timeout_sNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states that the tool blocks, polls server-side, and terminates on done/failed/cancelled or timeout, which are key behaviors. It does not, however, describe what happens after timeout (e.g., returns a status or throws an error) or the exact return value, leaving a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with two sentences that front-load the core action ('Block until...') and then add the polling detail and usage context. Every sentence earns its place; there is no fluff or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple wait tool with only two parameters and no output schema, the description covers purpose, behavior, and usage. However, it omits what happens on timeout (return vs. error) and does not describe the return value format, which is a meaningful gap given that the output schema is absent. The description is adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for missing parameter explanations. The description does not clarify job_id (where it comes from) or timeout_s beyond the default. While the parameter names are self-explanatory, the description adds no additional meaning, and the lack of guidance on the source or format of job_id is a notable omission.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action: it blocks until an index job reaches done/failed/cancelled or a timeout. The verb 'block' and the resource 'index job' are specific, and the terminal states are enumerated. It distinguishes itself from siblings like get_index_job, which would check status without blocking, so the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly provides a usage condition: 'Use this when you want a synchronous index call.' This tells the agent when to select this tool over the asynchronous alternatives. However, it does not name specific alternatives or explicitly state when not to use it, so it falls short of a perfect score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 20 tool updatesv2026.9.23
    • First observedexplain_code
    • First observedfind_callers
    • First observedfind_definition
    • First observedfind_references
    • First observedget_index_job
    • First observedgraph_explore
    • First observedhydrate_chunks
    • First observedindex_directory
    • First observedindex_file
    • First observedindex_graph
    • First observedindex_repo
    • First observedlist_index_jobs
    • First observedlist_indexed_sources
    • First observedlist_job_errors
    • First observedrebuild_all_graphs
    • First observedremove_indexed_source
    • First observedsearch_code
    • First observedsearch_code_enhanced
    • First observedsource_staleness
    • First observedwait_for_index_job

TDQS

A3.8/5.0

Scored across 20 tools

Disambiguation4/5

Most tools have clearly distinct roles: indexing scopes, job tracking, search, and graph traversal are easy to separate. The main ambiguity is between search_code, search_code_enhanced, and explain_code, and between find_callers and find_references, but the descriptions give enough usage guidance to make the intended choice.

Naming Consistency4/5

The set mostly follows a clean verb_noun snake_case pattern: index_*, get_*, list_*, find_*. Minor deviations like 'source_staleness' (a noun phrase instead of a verb), 'search_code_enhanced' (suffix modifier), and 'rebuild_all_graphs' (bulkier form) keep it from a perfect 5.

Tool Count3/5

20 tools is on the heavy side, and several are near-variants of the same operation: three search entry points and two graph-build tools add surface area. However, the server covers distinct responsibilities—indexing, job management, search, and graph analysis—so the count is not unreasonable.

Completeness4/5

The surface covers the full indexing lifecycle—create, list, track, force-reindex, and remove—along with code search and graph navigation. Minor gaps remain, such as no dedicated source-detail tool, no scoping for find_callers, and no explicit incremental update beyond force reindexing.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables AI agents to query and analyze code across multiple repositories through a unified knowledge graph, with tools for symbol search, impact analysis, and graph algorithms.
    31 npm
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    AST-aware semantic code search engine for AI agents, enabling code retrieval by intent with call-graph context and optional LLM enrichment.
    1
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    Enables coding agents to navigate and query source code by providing context, symbols, and call graph information through a graph index.
    4
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Provides AI agents with causal code memory by indexing repositories into a graph of symbols and edges, enabling context-aware retrieval of relevant code slices.
    23 PyPI
    3
    MIT