vex
Vex
Fast hybrid structural + semantic code search. Vector + index.
Why Vex? · How It Compares · Installation · Quick Start · Commands · Configuration · How Search Works · Benchmarks · Supported Languages · Integration · Testing · Architecture
$ vex check "TelemetryProcessor" # 4ms — does it exist? where? (exact name)
$ vex show "TelemetryProcessor" # extract the class body (not the whole file)
$ vex usages "Config" --strict # who references this symbol? (binder-resolved, no noise)
$ vex callers "process_event" # who calls this function? (~4ms; covers module-scope + Python/Java decorators)
$ vex implementations "BaseService" # who extends/implements this?
$ vex search "timeout retry" # fuzzy / multi-word — BM25 finds rare body terms
$ vex search "handle alert" --semantic # find by meaning, not just name
$ vex pattern 'fn $NAME($$$) -> Result' --lang rust # AST pattern matching (like ast-grep)
$ vex similar "PaymentService" # semantically close symbols
$ vex duplicates --threshold 0.95 # near-duplicate pairs
$ vex bundle --mode symbol --symbol Foo # body + callers + callees + similar in 1 callPick the right tool: vex check for "does Foo exist?", vex search for "find me something about retries". search is a ranked blend — it surfaces neighbors (callers / imports) when no symbol literally matches, which is great for exploration and wrong for exact-name lookup. v1.15.0 prints a stderr hint when an identifier-shaped search returns 0 FST hits.
Why Vex?
~4-5ms search after indexing — FST-based O(query_len) lookup, not O(symbols); constant regardless of project size. Requires a pre-built index. Indexing is a one-time cost (hundreds of ms on typical projects) and builds more than a plain text index — FST + BM25 + call graph + type-hierarchy + trigram skip-index — so it trades a slower build for far cheaper, richer queries (see Benchmarks)
3-channel hybrid search — structural FST (names) + BM25 (rare body terms) + semantic HNSW (meaning), fused via Reciprocal Rank Fusion. Find symbols when you don't know the exact name AND when generic semantic-only search would be too noisy
Persistent call graph —
vex callers/vex calleesread from a persistent index built at index time (~4ms), not a live tree-sitter scan (seconds):callersis a name-keyed FST,calleesis a dense CSR index (v9+). Module-scope expressions are reported via synthetic<module:path>callers (Phase 14.1); Python + Java function/method decorators (Phase 14.2), Kotlin annotations + C# method/constructor attributes (Phase 14.2.2), and TypeScript method decorators + Rust outer attributes (Phase 14.2.1) emit forward edges to their targets. Class-level decorators remain invisible — seedocs/LIMITATIONS.mdPluggable embedder —
Embeddertrait + registry; swap MiniLM-L6-v2 for future code-specific models (BGE, CodeBERT) without touching call sitesToken-efficient — compact output saves typically 6x fewer tokens than grep on average lookups (up to 217x on minified JS/CSS);
vex showextracts just the symbol body instead of the whole file19 languages indexed via tree-sitter, with three coverage tiers: type-aware
--strict usageson 8 binder languages (Rust / TypeScript / Python / C# / C++ / Go / Java / Kotlin); indexed pattern prefilter on 15 T1+T2a languages; baseline structural + semantic search on all 19 (see Supported Languages for the matrix)Single binary, zero config — no LSP servers, no databases, no Docker. Just
vex index && vex check Foo
Related MCP server: better-code-review-graph
What Vex isn't
vex is a static-analysis indexing tool, not a language server. Set expectations honestly:
Not an LSP replacement. No go-to-definition into third-party packages, no rename refactoring, no type-checking, no hover docs. For those, keep your LSP.
vex searchis a ranked blend, not an exact-name lookup. Structural FST + BM25 + semantic fused via RRF return relevance-ordered results — when no symbol literally namedFoolives in the index (imported from a dependency, deleted, typo), BM25 may surface callers / imports as if they were the definition. For exact-symbol questions ("does it exist?", "show me the body", "who calls it?") usevex check Foo/vex show Foo/vex usages Foo --strict— they bypass the ranker. v1.15.0 prints a one-line stderr hint when an identifier-shaped query gets zero FST hits.No dynamic-dispatch visibility. Decorator routing (
@router.get("/path")), string-resolved factories (uvicorn.run("main:app")), reflection (getattr(obj, name)()), and macro-expanded references are all invisible to every vex command.vex grep '\bname\b'is the textual escape hatch.vex callershas uneven coverage outside function scope. Module-level expressions likeapp = create_app()are reported via synthetic<module:path>callers (Phase 14.1). Python + Java function/method decorators (Phase 14.2), Kotlin annotations + C# method/constructor attributes (Phase 14.2.2), and TypeScript method decorators + Rust outer attributes on fns/methods (Phase 14.2.1) emit forward edges —vex callers GetMappinglists every Spring handler,vex callers HttpGetevery ASP.NET action,vex callers testevery#[tokio::test]. Class-level decorators (14.6) remain on the roadmap.vex usagesquality varies by language. 8 binder-supported languages get refactor-grade--strictrefs; the other 11 use an identifier scanner with a higher false-positive rate.
See docs/LIMITATIONS.md for the full coverage matrix, concrete repros, and recommended workarounds per query type. Read it before evaluating vex on a Python/FastAPI/Django codebase — the framework patterns are the most-flagged gaps.
How It Compares
vex | ripgrep | ast-index | ast-grep | Serena | |
What it searches | Symbol definitions | All text | Symbol definitions | AST patterns | Symbols (via LSP) |
Requires indexing? | Yes (~0.3-1s) | No | Yes (faster build) | No | No |
Search speed | ~4-5ms (pre-built FST, constant) | scales w/ corpus (~8ms small → 100ms+ large) | ~8-12ms (SQLite) | ~30ms (scan) | LSP-dependent |
Semantic search | HNSW + embeddings | -- | -- | -- | -- |
Pattern matching |
| regex only | -- |
| regex only |
Index size | ~1.5-2x smaller than ast-index | no index | SQLite + FTS5 | no index | no index |
Token efficiency | 6-217x fewer than rg | baseline | ~3x fewer than rg | N/A | N/A |
Symbol body extraction |
| -- | -- | -- | -- |
Languages | 19 | any | 10+ | 10+ | 40+ (LSP) |
Refactoring | -- | -- | -- | -- | rename, move, inline |
Runtime deps | none | none | none | none | Python + LSP |
Note: vex search speed assumes a pre-built index. Ripgrep and ast-grep require no upfront indexing and work immediately on any directory. The tradeoff is amortized: if you search the same codebase many times (typical in agent workflows), the one-time indexing cost pays for itself.
Best for: fast symbol search in AI agent workflows where token efficiency matters. Not a replacement for LSP-based tools (no refactoring, no go-to-definition in dependencies).
Installation
# Homebrew (macOS/Linux)
brew tap tenatarika/tap
brew install vex
# crates.io (compiles from source; the crate is `vex-search`, the binary is `vex`)
cargo install vex-search --locked # → ~/.cargo/bin/vex
cargo install vex-search-mcp --locked # → ~/.cargo/bin/vex-mcp (MCP server)
# From source (any platform with a Rust toolchain)
git clone https://github.com/tenatarika/vex.git
cd vex
cargo build --release
cp target/release/vex ~/.local/bin/What cargo install vex-search (and any source build) needs:
Network at build time: the build downloads a prebuilt ONNX Runtime. Prebuilts exist only for
aarch64-apple-darwin,x86_64/aarch64-unknown-linux-gnuandx86_64/aarch64-pc-windows-msvc; on any other target (Intel macOS, musl, …) pointORT_LIB_LOCATIONat a local ONNX Runtime build.A C/C++ toolchain (Xcode Command Line Tools,
build-essential, or MSVC Build Tools) for the tree-sitter grammars. GCC must be 12 or newer: GCC 11 (the default on Ubuntu 22.04) can't compile the bundled numkong vector kernels. Installgcc-12 g++-12and build withCC=gcc-12 CXX=g++-12 cargo install vex-search --locked.Linux:
libssl-devandpkg-config(Fedora:openssl-devel). The HTTP stack links OpenSSL throughnative-tls.The first
vex index --semanticdownloads the ~86 MB embedding model; structural search needs no download.vex-search-mcponly installs the MCP server. It runs thevexCLI, so installvex-searchtoo and keepvexonPATH, or setVEX_BINto its full path.
Linux
Pre-built vex ships in every GitHub Release for x86_64-unknown-linux-gnu. It needs glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+); on older systems build from source:
curl -L https://github.com/tenatarika/vex/releases/latest/download/vex-x86_64-unknown-linux-gnu.tar.gz | tar -xz
mv vex ~/.local/bin/ # or: sudo mv vex /usr/local/bin/
vex --versionBuilt on the ubuntu-22.04 GitHub runner (glibc-linked). For older glibc distros, musl-based distros (Alpine, NixOS without nix-ld), or aarch64 Linux (Graviton, Pi 5, Ampere) — build from source via cargo build --release.
Windows
Pre-built vex.exe ships in every GitHub Release.
Download
vex-x86_64-pc-windows-msvc.tar.gzfrom the latest releaseExtract
vex.exesomewhere stable (e.g.C:\Users\<you>\bin\) —tar -xzf vex-x86_64-pc-windows-msvc.tar.gzfrom a recent PowerShell, or 7-Zip / WinRAR via right-click. Security note:vex.exeloads the bundledDirectML.dllfrom its own folder, so on a multi-user machine or shared drive prefer a directory other users can't write to (e.g.C:\Program Files\vex\— with the trade-off thatvex self-updatethen needs an elevated shell). See GPU_SUPPORT.md §6.Add that folder to
PATH(System Properties → Environment Variables → editPath→ add the folder)Open a fresh terminal and run
vex --version
To update, run vex self-update — it fetches the latest release, picks the right archive for your platform, verifies its signature, and replaces the binary in-place. On Windows it also installs/refreshes the bundled DirectML.dll sidecar (skipped when byte-identical; re-installed if an older self-update dropped it — updaters up to v1.16.0 extracted only the binary). Same command works on macOS and Linux too.
GPU acceleration is built into the prebuilt binaries — Windows ships with DirectML (any DX12 GPU, driver-only; the redist
DirectML.dllis bundled in the archive) and macOS arm64 with CoreML. NVIDIA CUDA is a source-build opt-in. Runvex gputo check, and see GPU Acceleration.
Quick Start
# Index a project (structural only — fast)
vex index --path /path/to/project
# Index with semantic embeddings (slower first time, downloads 86 MB model)
vex index --path /path/to/project --semantic
# Exact-name lookup (does this symbol exist?)
vex check "PaymentService"
# Extract a symbol's body (no whole-file read)
vex show "PaymentService"
# Fuzzy / multi-word search (returns ranked neighbors when no symbol matches)
vex search "payment processing" --semantic
# Find all usages of a symbol (--strict drops string-literal / comment / wrong-scope noise)
vex usages "IndexReader" --strict
# File structure outline
vex outline src/main.rs
# Find implementations of a trait/interface
vex implementations "Iterator"
# Callgraph: who calls / is called by a function (fast path via persistent index)
vex callers "process_event"
vex callees "process_event"
# Multi-hop call graph (v1.7)
vex paths "main" "process_event" # all caller chains from main → process_event
vex reachable "process_event" # everything that transitively reaches it
vex tests-for "process_event" # tests covering process_event (path globs + name heuristic; framework label per row)
# Symbol-level diff against a branch (v1.7)
vex diff --base main # what symbols did this branch change?
# Historical view of a symbol — every commit that touched it (v1.15.0; v1.16.0 expanded)
vex index --history # build the persistent history sidecar once
vex history "PaymentService" # ~10ms — every version reachable from HEAD
vex history "PaymentService" --diff # unified diffs between consecutive versions
vex history "Foo" --since 2026-01-01 --author alice --kind function
vex history "deleted_symbol" --exact-presence # exact commit set where each blob lived (revert-aware)
# Semantic similarity by existing symbol — explain what's actually similar (v1.7)
vex similar "PaymentService" --limit 5 --min-score 0.7 --explain
# Near-duplicate pairs with reasoning (v1.7)
vex duplicates --min-score 0.95 --min-body-lines 5 --explain
# Search with per-call scope + metadata filters (v1.7)
vex search "Repository" --include 'src/**' --exclude '**/*.gen.*' --visibility public --async-only
# Why did the search return these results? (v1.7)
vex search "Foo" --why 2>trace.json
# Bundle: 4 round-trips → 1 envelope (v1.9, Phase 13.2)
vex bundle --mode symbol --symbol PaymentService # body + callers + callees + similar
vex bundle --mode pr-impact --base origin/main # changed symbols + transitive callers + tests
vex bundle --mode project --top-n 30 # top-N by reverse call-graph indegree
# Diff-context filters on every search-shaped command (v1.9, Phase 13.7-D3)
vex search "Repository" --since-branched # only files changed since branching from main
vex usages "Config" --since HEAD~3 # refs within the last 3 commits
vex callers "Foo" --changed-only # working-tree changes only
# Extract just a symbol's body — replaces Read for a specific function/class
vex show "PaymentService" # full body of the class / fn
vex show "Foo" "Bar" "Baz" # multiple symbols in one call
# Smart show truncation for token efficiency (v1.9, Phase 13.3)
vex show "BigClass" --signature-only # just the signature line
vex show "PaymentService" --head 20 # first 20 lines of the body
vex show "Foo" --no-body # signature + docstring, no body
# Ranking-eval harness — CI regression guard (v1.9, Phase 13.12)
vex eval --bench benches/ranking_golden/queries.toml # nDCG@10 / recall@10 / MRR per query
vex eval --min-ndcg 0.85 # fail if mean nDCG drops below threshold
# Capability discovery for MCP clients (v1.9, Phase 13.0)
vex capabilities # JSON: protocol_version, signals, bundle_modes, …
# Fast existence check
vex check "Foo" "Bar" "Baz"
# Incremental update (re-parses only changed files, reuses unchanged from index)
vex update
# Watch mode (re-indexes on file changes)
vex watch
# Multi-repo: treat a set of sibling repos as one workspace (v1.22.0)
vex index --workspace # build every member of .vex-workspace.toml
vex search "RetryPolicy" --workspace # fan out, results grouped by repo
vex usages Config --strict --workspace # cross-repo strict refs (v7+ index)
vex watch --workspace # keep every member incrementally fresh
# Show index stats
vex status
# GPU doctor — is the compiled EP actually engaging on this machine? (v1.16.0)
vex gpu # probes the compiled-in EP with strict registration
vex gpu cuda # narrow to one EP
vex gpu --enable # persist working device to VEX_DEVICE
# Shell completions
vex completions zsh > ~/.zfunc/_vexCommands
Command | Description |
| Build full index. |
| Hybrid search: structural + BM25 + semantic (when |
| Extract symbol body from source (saves tokens vs full file read). Same metadata + kind filters as |
| Find symbols semantically close to an existing one (HNSW nearest neighbors). |
| List near-duplicate symbol pairs by embedding similarity. |
| Find all references/usages of a symbol. Non-strict path = FST lookup; v1.20.0 strips the row at the symbol's own definition line and |
| One-call delete-safety blast-radius report (since v1.20.0). Composes four reference channels — strict refs (binder-resolved), FST refs, |
| AST pattern matching with metavariables ( |
| Show file structure, optionally filter by symbol kind. |
| Find types that extend/implement a base class, trait, or interface (incl. generic-parameterised: |
| Transitive-down closure over |
| De-facto modules: clusters of symbols that call/reference each other, computed on |
| Direct callers of a function (fast path via persistent call graph; falls back to live tree-sitter scan when the index is missing). |
| Direct callees of a function (same fast path). |
| Enumerate all caller chains from |
| Transitive set of symbols whose callees reach |
| Test functions that transitively cover |
| Symbol-level diff between an arbitrary git revision and the working tree: added / removed / moved-within-file / body-changed entries. |
| Unified multi-source bundle (since v1.9) — replaces 4 round-trips ( |
| Fast existence check — which symbols exist in the index? |
| Regex content search (no index needed). |
| Incremental update — re-parse only changed files, reuse unchanged symbols from existing index. |
| Watch filesystem, auto re-index on changes. |
| Show index stats: symbol count, size, embeddings, call graph, BM25, GPU support. |
| Diagnose GPU acceleration: prints the execution provider compiled into this binary and actively probes whether it engages on this machine (a silent CPU fallback shows as |
| Generate shell completions (bash, zsh, fish). |
| Create a default |
| Manage |
| Print the machine-readable capability matrix (since v1.9): |
| Run the ranking-evaluation harness against a hand-curated golden query set (since v1.9); reports nDCG@10 / recall@10 / MRR per query and aggregated. CI regression guard — fails when mean nDCG drops below |
| NEW (v1.15.0); expanded in v1.16.0 (Phase 14.9). Every historical version of a symbol reachable from a chosen tip. With |
| Update vex to the latest GitHub release. Replaces the running binary in place. Works on Linux, macOS, and Windows. |
Per-query filters (every search-shaped command)
All search-shaped commands accept scope filters. Specific filters vary by command:
--include <glob>/--exclude <glob>(repeatable, gitignore syntax) — per-call path scoping that doesn't require re-indexing.--excludewins over--include. Example:vex search Foo --include 'src/**' --exclude '**/*.gen.*'.--exclude-tests(MCP:exclude_tests: true) drops test files (tests/,*_test.*,test_*.py,*.spec.ts,__tests__/,tests.rs, … — the same set asvex tests-for); it composes with--include/--exclude, is recorded in the--whytrace, and is applied by every command that takes the scope flags (search,usages,callers/callees,impact,grep,show,pattern,implementations/subtypes,similar/duplicates,modules,paths,reachable,diff,bundlepr-impact — there it filters the changed files and the transitive-caller/test rows;bundlesymbol/projectmodes ignore all scope filters).vex tests-forrejects--exclude-tests(exit 2): it lists test functions, so the flag would always return nothing. Path-based only: Rust unit tests inside a#[cfg(test)] mod testsblock of a non-test file are not excluded.--filter-path <substring>(alias--filter) — path-substring filter onsearch,show,usages,grep,similar,duplicates. Composes AND with the globs.
vex search / vex show additionally accept:
--visibility <public|private|protected|internal>— keep only symbols whose signature carries the explicit keyword. Defaults aren't inferred (bare Rustfn foo()does NOT match--visibility private).--async-only/--no-async— keep or exclude async / Kotlin-suspendsymbols.--static-only,--sealed-only— restrict to static class members or sealed (or Java-final) types.
Reasoning flags
vex search --whyprints a JSON trace to stderr (the result list stays on stdout):normalized_query, per-channel hit counts (FST / BM25 / semantic), fallbacks engaged (fuzzy), and the active filter snapshot.vex pattern --whyprints a JSONScanTraceto stderr after the result list:mode(indexed/live_scan),root_kind_inferred,candidate_files/total_files, andfallback_reasonwhen the indexed prefilter was skipped (no-index,no-skeleton-section,empty-section,grammar-drift,partial-section,index-open-error). MCP callers see the same JSON under_meta.why.vex similar --explain/vex duplicates --explainadd ajaccardoverlap score plus a truncated unified diff between the two bodies, so you can decide whether two semantically-clustered symbols are actually duplicates before acting.
Multi-repo workspaces (--workspace) — v1.22.0
Declare a set of sibling repos in a .vex-workspace.toml and run any command with --workspace to fan out across all of them, grouped by repo:
# .vex-workspace.toml (at the directory that contains the repos)
members = ["./api", "./worker", "./shared-lib"]vex index --workspace # build every member into its own per-repo index
vex update --workspace # incremental refresh, per-repo changed/deleted counts
vex usages Config --strict --workspace # cross-repo strict refs, grouped by repo--workspaceis accepted byindex,update,search,grep,check,usages,impact,callers,callees,reachable,modules, andwatch. Each member keeps its own.vex.toml(excludes / embedder / sections / cache).Reference and call-graph resolution is per-repo by default. The one exception is
vex usages <name> --strict --workspace, which resolves a reference in repo B to a symbol defined in repo A via a gtags-style name fallback (rendered as aname-resolvedsub-tier). Requires v7+ index — re-runvex indexafter upgrading.A member missing a capability (
--stricton an old index, a call graph forreachable) is reported unavailable for that repo instead of aborting the whole fan-out.--workspaceconflicts with--why. Seedocs/MULTIREPO.mdand LIMITATIONS §7.
Configuration
Create a .vex.toml in your project root to customize vex behavior:
vex init # generates .vex.toml with commented defaults# .vex.toml
# Glob patterns to exclude from indexing (gitignore syntax, on top of .gitignore)
exclude = [
"vendor/**",
"node_modules/**",
"*.generated.go",
]
# Output format — "compact" (default since v1.10.1; single-line records),
# "text" (verbose multi-line), or "json" (envelope for MCP / tools).
# format = "text"
# Enable semantic embeddings by default
semantic = true
# Automatically update index before search if stale
# auto_update = false
# With auto_update, refresh in the background rather than blocking the query:
# answers come from the index on disk, the response is flagged stale, and the
# rebuild lands for the next query. Trades freshness for latency.
# async_update = false
# GPU device for semantic indexing (GPU-enabled builds only). "auto" uses the
# compiled-in GPU EP when it initializes, else CPU; or "cpu"/"cuda"/"directml"/
# "coreml". `gpu = true/false` is shorthand for auto/cpu. See GPU Acceleration.
# device = "auto"
# gpu = true
# Embedder model: minilm-l6-v2 (default), jina-code, bge-base-en-v1.5,
# bge-large-en-v1.5, mxbai-large. Changing it requires a reindex.
# Set globally across projects with the VEX_EMBEDDER env var (this file wins).
# embedder = "minilm-l6-v2"
# VCS backend for diff-scoping (--since/--since-branched/--changed-only).
# "auto" (default) detects .git/.svn/.arc; "git" | "none" | "arc" | "svn".
# git, arc (Yandex Arc), and svn (Subversion) are all functional backends;
# svn declines --since-branched (no merge-base). "none" disables diff-scoping.
# Overridden by the --vcs flag and the $VEX_VCS env var. See docs/VCS-BACKENDS.md.
# vcs = "auto"CLI flags always override config values. Use --no-semantic to explicitly disable semantic mode when the config enables it. The VEX_DEVICE and VEX_EMBEDDER environment variables act as global defaults across all projects (lowest precedence, below .vex.toml) — see GPU Acceleration.
Keeping config out of the repo
Don't want a .vex.toml inside the repository (can't .gitignore it, shared checkout, etc.)? vex never creates one on its own — only vex init writes it — and you have two ways to keep config external:
--config <path>/$VEX_CONFIG— point vex at a config file anywhere on disk. It replaces the in-repo lookup entirely, so the repo stays clean:vex --config ~/vex/this-repo.toml search Foo export VEX_CONFIG=~/vex/this-repo.toml # or set it once per shell--configbeats$VEX_CONFIG; a missing/invalid path is a hard error (vex won't silently fall back). Relative paths inside that file resolve against the file's own directory.A parent directory — config lookup walks up from the project to the filesystem root, so a
.vex.tomlplaced in any ancestor (e.g.~/work/.vex.toml, or~/.vex.tomlfor a machine-wide default) is picked up for every repo beneath it, with none living in the repos themselves.
The index itself is never written into the repo — it lives in the cache dir (--cache-dir / $VEX_CACHE_DIR / platform cache), so a clean repo is just a matter of config placement.
Staleness Detection
Vex detects when the index is stale and warns before search:
$ vex search "Config"
Warning: index may be stale (HEAD changed). Run `vex update`.How it works: on every search, vex compares the git HEAD stored at index time with the current HEAD (~0.1ms, single git rev-parse). If HEAD changed → stale. For non-git repos, falls back to mtime comparison — and since v1.11 (H11), when mtime fires, vex streams a xxh3_64 content hash of the file and compares it to the manifest. If the hash matches, the touch was cosmetic (git checkout, rustfmt no-op, rsync --times) and the file stays Fresh; only a real content change re-triggers indexing.
Auto-update: skip the warning and update inline:
# Per-command
vex search "Config" --auto-update
# Always (in .vex.toml)
auto_update = true
# Disable staleness check entirely
vex search "Config" --no-stale-checkGPU Acceleration
Semantic indexing (--semantic) can run the embedding model on a GPU — a large win on a full/cold index. On an RTX 3080 over a 28k-symbol C++ module, embedding the default MiniLM model was 51× faster on CUDA and 29× on DirectML vs CPU (full benchmark + design notes in docs/GPU_SUPPORT.md).
Two layers — the binary, and the device:
Prebuilt binaries bake in a driver-only GPU EP: Windows → DirectML (any DX12 GPU — NVIDIA/AMD/Intel; the redist
DirectML.dllis bundled in the archive), macOS arm64 → CoreML. No SDK, no extra install. The Linux prebuilt is CPU-only.CUDA is a source-build opt-in (fastest on NVIDIA — ~1.75× DirectML):
cargo install --git https://github.com/tenatarika/vex vex-search --features gpu-cuda. Needs the CUDA 12 runtime + cuDNN 9 onPATH(the NVIDIA driver alone is not enough — it ships onlynvcuda.dll, not the runtime/cuDNN). Source builds for the others:--features gpu-coreml/gpu-directml.
Selecting the device (vex index / vex update):
--gpu/--no-gpu— force GPU (Auto) or CPU.--device cpu|auto|cuda|directml|coreml— pick a specific execution provider.Default is Auto: use the compiled-in EP when it initializes, else silently fall back to CPU. On a CUDA-enabled binary Auto prefers CUDA → DirectML → CoreML; on the standard Windows prebuilt only DirectML is compiled in, so Auto uses DirectML regardless of whether the PC could do CUDA.
A tiny incremental
vex updatestays on CPU (the GPU warm-up isn't worth a handful of symbols); cold/large--semanticbuilds use the GPU.
Is the GPU actually being used? Run vex gpu — it reports the compiled EP and actively probes it, so a silent CPU fallback shows as FAILED with targeted setup remediation. vex gpu cuda probes a single EP; vex gpu --enable persists the working device to VEX_DEVICE (user env via setx on Windows; prints the export line to add on macOS/Linux). A stale VEX_DEVICE pinned to a GPU EP that a later (e.g. CPU-only) build lacks degrades to CPU rather than erroring.
Environment variables
Variable | Effect |
| Global default device ( |
| Global default embedder id (e.g. |
| Turn ORT's silent CPU fallback into a hard error — proves whether the GPU engaged. ( |
| Advanced: hard cap on the GPU arena VRAM. Set it generously (≥ working set) or it OOMs on long-context batches. |
| Advanced: tune length-aware batch sizing (the |
Output Formats
# Compact single-line records — default since v1.10.1 (token-efficient, agent-friendly)
vex search "Foo"
# Verbose multi-line / human-readable
vex search "Foo" --format text
# JSON envelope (for MCP / tool integration; what `vex-mcp` parses)
vex search "Foo" --format jsonPin a different default in .vex.toml via format = "text" if you want the verbose multi-line view at the terminal.
JSON envelope (v1.11.0 — BREAKING for bare-array parsers)
Every --format json subcommand wraps its payload in the Phase 13
envelope. Single shape, easy to detect via protocol_version:
{
"protocol_version": "v1",
"capabilities": { /* see `vex capabilities` */ },
"_meta": { "vex.dev/index_age_ms": 1200, "ttlMs": 30000, "cacheScope": "project" },
"results": [ /* the actual data, shape depends on the subcommand */ ]
}Pre-v1.11 only search and bundle returned this envelope; the other
subcommands (show, usages, pattern, grep, implementations,
callers, callees, paths, reachable, tests-for, check,
similar, duplicates, diff, outline, index, update,
status, eval) emitted bare arrays / objects. Migration: pre-1.11 jq '.[0].name'
or data[0]['name'] now needs jq '.results[0].name' /
data['results'][0]['name']. Detect the envelope via
response.get('protocol_version') == 'v1' to support both shapes
during a rollout window.
How Search Works
Structural Search (default)
Searches by symbol name using an inverted index with CamelCase splitting:
"PaymentService"— exact match"Payment"— prefix match, finds PaymentService, PaymentGateway"payment"— case-insensitive, also finds via CamelCase tokens
Semantic Search (--semantic)
Embeds your query with MiniLM-L6-v2 (384-dim vectors) and finds symbols with similar meaning:
"parse source code files"findsparse_file,extract_refs,parse_file_symbols"database storage"findspopulate_db,create_10k_db,add_root_persists_to_db"find implementations of an interface"findsfind_implementations,test_interface_extends
BM25 Channel (auto-on when index has BM25 data)
A classic Okapi BM25 (K1=1.2, B=0.75) over symbol body tokens — identifiers, signatures, docstrings. Closes the gap between "exact name" (structural) and "general meaning" (semantic): finds rare body terms like timeout, retry, singlestore, idempotency_key that aren't part of any symbol name. Since v1.11 (Phase 8.4) body tokens are also extracted from TOML / YAML / HTML / CSS values, so vex search "production endpoint" --semantic can hit a [server] table with endpoint = "https://...". Pass --no-bm25 to disable per-call.
Hybrid Search (3-way RRF)
When the index has all three channels (built with --semantic), vex search fuses structural + BM25 + semantic using Reciprocal Rank Fusion. Symbols hit by ≥2 channels rank as Hybrid; symbols unique to one keep their original match type. Cuts both structural-noise and semantic-blur in the same query.
Usages (FST)
References stored in an FST (Finite State Transducer) — zero-copy lookup from mmap with prefix search support.
Symbol clusters (vex modules)
vex index groups symbols into clusters of code that call or reference each other (deterministic Leiden-CPM over the call, reference and hierarchy edges; no randomness, so two indexes of the same tree agree). vex modules reads them back:
vex modules # clusters of >= 3 symbols, largest first
vex modules --members 5 --sort cohesion --include 'src/**'
vex modules IndexReader # the cluster of one symbol, with its members
vex modules --format text # text output (shown below)Modules — leiden-cpm/1 γ=1/8 · 475 clusters (≥3, showing 3) · 1,469 unclustered · 1,464 not eligible
#60 src/cli/ 39 symbols cohesion 0.49 hubs: OutputFormat, print_envelope, default_meta_for
#23 crates/vex-mcp/src/tools/ 38 symbols cohesion 0.68 hubs: opt_bool, build_command, opt_u64
#286 src/pattern/matcher/tests.rs 32 symbols cohesion 0.78 hubs: parse_pattern, find_matches, Segment(Output above is from this repository at the time of writing; cluster ids are ordinals in the section, not ranks.)
A cluster's
labelis the deepest path prefix holding at least 60 % of its members' files.cohesionisinternal / (internal + cut)edge weight; the hubs are the three members best connected inside the cluster.--include/--exclude/--exclude-testsfilter members and hubs (out-of-scope hubs are dropped); a cluster is shown when at least one member is in scope, and itssizeis the in-scope count (size_at_buildkeeps the build-time count in JSON).Symbol mode reports a per-match
status:clustered,unclustered(isolated),not_eligible(headings, modules, markup/config languages) ornew_since_build.Clusters are computed by
vex indexand carried, frozen, acrossvex update. After an update the JSON hasstale: trueandnew_since_build, and text output ends with a!line; runvex indexto recompute. This is separate from_meta.vex.dev/stale, which still means "index older than the working tree".--limitmust be at least 1; with a SYMBOL it caps the matched symbols (JSONsymbols_totalreports the uncapped count) and--min-sizeis ignored. Exit codes:0with results,1when empty (the reason is inresults.empty_reasonand on stderr),2for a corrupt cluster section. With--workspace, clusters are per repo and--limitapplies per repo.
Type-aware refs (--strict)
vex usages --strict <name> reads the v5 reference_edges section
written by an LSP-style scope binder. For the languages with a
binder (Rust, TypeScript, Python, C#, C++, Go, Java, Kotlin) every ref is resolved at
index time against an in-file scope chain plus an import/use graph,
then serialised against the global symbol the user actually meant —
not just any line that mentions the spelling.
What this changes for the user:
Identifiers inside comments, doc-strings, string literals, and regex bodies are dropped (this filter is on for everyone, not just
--strict).A name shadowed by a
let/const/ fn param resolves to the inner scope, not the outer.A
use ext::Foo;/import { Foo } from './ext'/from ext import Foomakes a ref toFooresolve cross-file to whatever defines it in the index. For C++, quoted#include "..."(v1.14+) walks the transitive include graph via BFS to resolveFooagainst symbols defined in any reachable header. System headers<vector>/<string>and macro includes (#include MY_HEADER) stay unresolved by design.A name imported but never defined in the index stays
Unresolvedand produces no edge — better than a coincidental match.
Without --strict vex usages still works for every supported
language via the legacy refs FST; --strict simply trades recall
breadth for precision on the eight binder languages. v3 / v4 indexes
predating the binder bail with a "re-run vex index" message.
Structural Patterns (vex pattern)
Match code by shape rather than text. Works on every language vex
parses; an indexed prefilter (via the v6 pattern_skeletons section)
speeds up candidate selection on the 15 languages that emit skeletons
(every language except Bash, Lua, YAML, and TOML, which live-scan).
Syntax:
$NAME— capture a single identifier or balanced expression. Same name appearing twice enforces a back-reference:record($X, $X)matchesrecord(state, state)and rejectsrecord(state, other).$_— wildcard (matches without capturing).$$$— anonymous ellipsis (matches anything up to the next literal; spans newlines).$$$BODY/$$ARGS— named multi-line ellipsis. Functionally identical to$$$but captures the consumed text under the given name;$$$BODYreads naturally for block bodies,$$ARGSfor parameter lists. Back-reference equality also applies.&&(space-flanked) — AND composition. Both sub-patterns must match in the same file, and shared metavar names must capture the same text in both:struct $S && impl $Smatches files that have both shapes for the same$S.||(space-flanked) — OR composition (union, deduped by(path, line)).&&binds tighter than||.Composition operators only fire at bracket / quote depth 0, so
record($X, $X)andf($X && $Y)stay single patterns.
Indexed prefilter: when a v6 index is present, the leading literal
keyword of the pattern (fn, struct, class, def, impl, …) is
mapped to a tree-sitter node kind, and vex pattern walks only the
files whose persisted skeletons contain that kind. Visibility / async
/ export modifiers in front of the keyword are stripped before the
match (pub async fn $F infers function_item correctly). Falls
back to live-scan on grammar drift, missing section, or a partial
section after vex update — --why reports the exact reason.
Examples:
# Multi-line function body with named captures
vex pattern 'fn $NAME($$ARGS) -> Result<$T, $E> { $$$BODY }' --lang rust
# Both struct and impl for the same type in one file
vex pattern 'struct $S && impl $S' --lang rust
# Interface OR class with the same name
vex pattern 'interface $N || class $N' --lang typescript
# See which mode and what narrowing happened
vex pattern 'fn $N($$$)' --lang rust --why 2>trace.jsonBenchmarks
Compared against ast-index v3.31.0 (SQLite + FTS5) and ripgrep 15.1.0.
Methodology: indexing re-measured 2026-08-12 on Apple Silicon (macOS), vex v1.25.5, release build, cold cache, non-semantic index. Search figures are unchanged from the 2026-07-11 / v1.25.1 run (nothing in v1.25.2-v1.25.5 touches the search path) and are the average of 10 runs. Reproduce with ./benches/bench.sh (point VEX_BENCH_LARGE_PROJECTS at your own repos for larger corpora). Numbers are machine-specific — treat the ratios, not the absolutes, as the signal.
Indexing
Project | vex | ast-index | vex size | ast-index size |
Small (vex itself, 6.8K symbols) | 442 ms | 254 ms | 4.6 MB | 6.7 MB |
Medium (ast-index repo, 2.3K symbols) | 204 ms | 109 ms | 1.6 MB | 3.4 MB |
Honest read: vex indexing is ~1.7-1.9x slower than ast-index — it builds far more at index time (FST + BM25 + persistent call graph + resolved reference edges + a v8 type-hierarchy section + a trigram skip-index + pattern skeletons + symbol clusters (v9)), where ast-index builds a SQLite + FTS5 store. That one-time cost buys the constant-time queries below; the resulting index is still ~1.4-2x smaller on disk (mmap + FST vs SQLite). These measurements are from v1.25.5 (before v9); v9 adds the cluster pass — re-measure pending. The gap was ~2.5-3x when last measured at v1.25.1; v1.25.5 is the main reason it has narrowed — a single-variable A/B against v1.25.4 put its cold-index gain at −34% and −31% on two corpora, from parsing each file once and sharing the tree across all extractors instead of re-parsing it per extractor. Projects indexed with --semantic are slower again (ONNX embedding generation) and produce a larger index.
Search: vex vs ast-index vs ripgrep
Medium project (ast-index repo, 31K lines Rust, avg 10 runs):
Query | vex | ast-index | rg -w | vex vs rg |
| 4.7 ms | 8.3 ms | 9.2 ms | 2.0x |
| 4.6 ms | 8.2 ms | 8.6 ms | 1.9x |
| 4.6 ms | 7.9 ms | 8.7 ms | 1.9x |
| 4.7 ms | 11.7 ms | 8.6 ms | 1.8x |
Key takeaway: vex search is constant ~4-5 ms (FST O(query_len)) regardless of project size — this is the win the slower index build pays for. The ripgrep comparison is not apples-to-apples: rg scans raw text with no index, so it scales with corpus size (single-digit ms on this 31K-line repo, 100 ms+ on large ones), while vex does a pre-built FST lookup. The durable advantage is amortized and qualitative: vex returns only symbol definitions (precise, token-efficient), while rg returns every text occurrence (noisy, expensive in LLM contexts).
Pattern Matching (vex only)
Medium project (ast-index repo, Rust):
Pattern | Time | Matches |
| 31 ms | 50 |
| 27 ms | 44 |
| 29 ms | 50 |
ast-index and ripgrep do not support AST pattern matching.
Semantic Search
Illustrative (semantic capability, not a latency benchmark) — queries where structural search returns 0 results but semantic finds relevant symbols:
Query | Structural | Semantic |
"parse source code files" | 0 | 19 |
"database storage" | 0 | 20 |
"find implementations of an interface" | 0 | 20 |
"file system directory walker" | 0 | 20 |
"handle errors and exceptions" | 0 | 20 |
HNSW vs Brute-Force (semantic vector search)
The latency figures below were measured on an earlier build and not re-run in the v1.25.1 pass — read them as the scaling shape (HNSW stays flat, brute-force grows linearly), not current absolutes.
Semantic search embeds the query via ONNX (~55ms) then searches stored vectors. HNSW (usearch) replaces brute-force O(N) scan with O(log N) approximate nearest neighbor search:
Symbols | Brute-force | HNSW | Speedup |
333 | ~3 ms | ~3 ms | 1x |
11K | ~8 ms | ~3 ms | 2.3x |
20K | ~11 ms | ~3 ms | 4x |
100K (projected) | ~55 ms | ~3 ms | ~18x |
HNSW stays constant ~3ms regardless of index size. Brute-force grows linearly. Total semantic search latency is dominated by ONNX embedding (~55ms), so end-to-end speedup is modest for small codebases but critical at scale.
Mode | Latency |
Structural only | ~4 ms |
Hybrid (structural + semantic) | ~58 ms (HNSW) / ~66 ms (brute-force) |
LLM Token Efficiency
When an AI agent searches code, the output goes directly into the context window. Grep-based tools return every text occurrence — including comments, strings, variable usage, and matches in minified files — consuming tokens without adding signal.
vex returns only symbol definitions in a compact one-line format, drastically reducing token consumption:
vex compact | rg (grep) | Reduction | |
7 symbol lookups (typical) | ~220 tokens | ~1,300 tokens | 6x |
Queries hitting minified JS/CSS | ~270 tokens | ~58,700 tokens | 217x |
Example — searching for a class name on a large project:
# rg: 20 matches across imports, usage sites, comments, tests (2,045 chars)
$ rg -w "PreAggregatedConfig" .
./models.py:3602:class PreAggregatedConfig(models.Model):
./models.py:3610: pre_aggregated_config = PreAggregatedConfig.objects.get(...)
./serializers.py:48:from .models import PreAggregatedConfig
./tests.py:12: config = PreAggregatedConfig(...)
... (16 more lines)
# vex: 1 definition (93 chars)
$ vex search "PreAggregatedConfig" --format compact
C PreAggregatedConfig models.py:3602 class PreAggregatedConfig(models.Model):For an agent making 10-20 code lookups per task, vex saves 5,000-20,000 tokens per session compared to grep — reducing cost and leaving more context window for reasoning.
Supported Languages
19 languages indexed via tree-sitter. The capability columns:
Binder — does
vex usages --strictresolve refs through an LSP-style scope chain (Phase 11.1)?cross-fileincludesuse/importresolution;in-fileresolves within a file but treats imports as unresolved. The remaining languages fall back to the line-based scanner used by plainvex usages.Patterns — does
vex patternget the v6 indexed prefilter (Phase 11.4)?indexedmeans a persisted skeleton section narrows candidate files at query time;live-scanmeans tree-sitter walks every lang-matching file on each query. All 19 languages work withvex patternsyntax ($NAME,$$$BODY,&&/||); the prefilter just speeds up discovery for the 15 languages that emit a skeleton section (every language except Bash, Lua, YAML, and TOML).
Language | Extensions | Symbols | Imports | Binder | Patterns |
Rust |
| functions, structs, enums, traits, impls, types, constants |
| cross-file | indexed |
TypeScript/JS |
| classes, interfaces, enums, functions, arrows, type aliases |
| cross-file | indexed |
Python |
| classes, functions (incl. async, decorated) |
| cross-file | indexed |
C# |
| classes, interfaces, structs, enums, methods, properties |
| cross-file | indexed |
C/C++ |
| classes, structs, functions, methods, templates, enums |
| cross-file (v1.14 BFS over quoted | indexed |
Go |
| functions, methods, structs, interfaces |
| cross-file | indexed |
Java |
| classes, interfaces, enums, methods, constructors |
| cross-file | indexed |
Kotlin |
| classes, interfaces, objects, functions, properties |
| cross-file | indexed |
Ruby |
| classes, modules, methods | — | — | indexed |
Swift |
| classes, structs, enums, actors, protocols, functions |
| — | indexed |
PHP |
| classes, interfaces, traits, methods, functions |
| — | indexed |
SQL |
| tables, views, functions, triggers, indexes, schemas, types, sequences |
| — | indexed |
Markdown |
| headings (section structure) | — | — | indexed |
Bash |
| functions | — | — | live-scan |
Lua |
| functions, local functions, tables |
| — | live-scan |
CSS |
| rules, selectors, | — | — | indexed |
HTML |
| custom elements (hyphenated tag names) | — | — | indexed |
YAML |
| top-level keys | — | — | live-scan |
TOML |
| bare keys, dotted keys, tables | — | — | live-scan |
See docs/SUPPORTED_LANGUAGES.md for grammar
versions, ABI level, and the runbook for adding a language or upgrading a
grammar. Adding a language to the indexed-Patterns tier is one
allowlist edit in src/pattern/skeleton/kinds.rs — the Phase 11.4
follow-up promotion (Go → Java → Kotlin → C# → C++ → Swift → PHP →
Ruby, plus SQL / Markdown / CSS / HTML) is complete; only Bash, Lua,
YAML, and TOML remain on live-scan.
Index Location
macOS: ~/Library/Caches/vex/<hash>/index.vex
Linux: $XDG_CACHE_HOME/vex/<hash>/index.vex (fallback: ~/.cache/vex/<hash>/index.vex)
Windows: %LOCALAPPDATA%\vex\<hash>\index.vex (fallback: %USERPROFILE%\AppData\Local\vex\<hash>\index.vex)Each project gets its own index based on a hash of the canonical project root path (xxh3). Overrides:
--cache-dir <path>— point vex at a custom cache directory$VEX_CACHE_DIR— environment variable (lower precedence than--cache-dir)cache_dirin.vex.toml— configuration file (lowest precedence)
Known limitations
vex is a static-analysis tool — some real call sites and references are invisible by construction. The headline gaps:
vex callersoutside function scope — Module-level expressions are reported via synthetic<module:path>callers (Phase 14.1). Python + Java function/method decorators (Phase 14.2), Kotlin annotations + C# method/constructor attributes (Phase 14.2.2), and TypeScript method decorators + Rust outer attributes on fns/methods (Phase 14.2.1) emit forward edges. Remaining gap: class-level decorators (Phase 14.6); Rust#[derive(...)]is intentionally filtered.vex usagesquality depends on language. Rust / TypeScript / Python / C# / C++ get--strict(binder-resolved refs from the v5reference_edgessection, Phase 11.1). Other languages use a line-based identifier scan with a higher false-positive rate.Dynamic dispatch is invisible. String-resolved factories (
uvicorn.run("main:app")), task queues (celery_task.delay()), reflection (getattr(obj, name)()) — none of these produce edges.Workaround:
vex grep '\bname\b'is the exhaustive textual fallback. Slower (~50 ms) but never misses a hit.
See docs/LIMITATIONS.md for the full coverage
matrix, repros, and recommendations per query type.
Troubleshooting
Surfacing internal warnings
Vex emits structured logs via the tracing crate at parse/store
boundaries — failed grammar loads, mmap reopens, manifest mismatches,
and so on. By default RUST_LOG is unset, so only the most critical
diagnostics make it to stderr.
When a search returns surprising results or an index command behaves oddly, raise the log level:
RUST_LOG=vex=warn vex search Foo
RUST_LOG=vex=info vex index # noisier — file-level progressFor what the search engine actually did (per-channel hit counts, fuzzy fallback engagement, applied filters), use the structured trace instead:
vex search Foo --why 2>trace.json # trace lands on stderr as JSONSee docs/MCP-SCHEMA.md for the --why /
why: true JSON shape.
Integration
Claude Code (CLI Integration)
The recommended way to integrate vex with Claude Code is via CLAUDE.md rules (see below). Vex runs as a CLI tool — Claude Code calls it directly via Bash, no MCP server needed.
Setup:
# Install vex
brew tap tenatarika/tap && brew install vex
# In your project
cd /path/to/project
vex init # create .vex.toml
vex index # build index (add --semantic for meaning-based search;
# add --history for `vex history <Symbol>` archaeology queries — v1.15.0/v1.16.0)Then add .vex.toml config for auto-update so Claude always searches a fresh index:
# .vex.toml
auto_update = true
# format = "compact" # already the default since v1.10.1 — set "text" if you'd rather see verbose outputMulti-repo (v1.22.0): if Claude Code is working across several repos at once, drop a .vex-workspace.toml at the common parent and tell Claude to add --workspace to its vex calls — e.g. vex usages Config --strict --workspace to trace a symbol's references across every repo, or vex check Foo --workspace to see which repos define it. Results come back grouped by repo. See Multi-repo workspaces.
Claude Code (MCP Server)
Alternatively, vex includes an MCP server (vex-mcp) that exposes all commands as MCP tools. Note: Homebrew installs only vex (not vex-mcp). Since v1.11.2 a prebuilt vex-mcp binary ships in every release alongside vex for the three triples the build matrix covers: aarch64-apple-darwin (macOS Apple Silicon), x86_64-unknown-linux-gnu (Linux), and x86_64-pc-windows-msvc (Windows). Intel-Mac and other triples still require the source build below.
Easiest setup (v1.15.0+):
vex mcp install --agent claude-codeThis runs claude mcp add --scope user --transport stdio vex --env VEX_ROOT=<root> -- <vex-mcp> for you (v1.27.1+). If the claude CLI is not on PATH, it prints that command instead and writes nothing. Releases before v1.27.1 wrote ~/.claude/claude_desktop_config.json, which Claude Code does not read; re-run the command after upgrading.
Manual setup:
# 1. Download the prebuilt for your platform from
# https://github.com/tenatarika/vex/releases/latest
# e.g. vex-mcp-aarch64-apple-darwin.tar.gz / vex-mcp-x86_64-pc-windows-msvc.tar.gz
# 2. Extract and put the binary on PATH (or remember the full path).
# Source build (if you prefer or are on an unsupported triple)
cargo build --release -p vex-search-mcp
# Register with Claude Code (user scope; Claude Code keeps it in ~/.claude.json)
claude mcp add --scope user --transport stdio vex \
--env VEX_ROOT=/path/to/your/project --env VEX_DEVICE=auto \
-- /path/to/vex-mcp
# Or project scope: commit a .mcp.json at the project root
# (see integrations/claude-code/mcp.json)VEX_DEVICE (v1.16.0) picks the GPU execution provider when the binary was built with gpu-cuda / gpu-directml / gpu-coreml — relevant when an MCP-driven index / update call rebuilds semantic embeddings on a large repo (51× CUDA / 29× DirectML over CPU on MiniLM-L6). auto is safe on CPU-only builds (degrades silently). Run vex gpu once to confirm the EP actually engages.
MCP Tools (28):
search— 3-way hybrid (structural + BM25 + semantic); acceptsfilter/include/exclude/kind/context_path/no_bm25/--why/ metadata filters / diff-scope (since/since_branched/changed_only)find_symbol— exact name lookupfind_similar— semantic search by free-form descriptionsimilar— nearest neighbors of an existing symbol (explainadds Jaccard + diff); diff-scopeduplicates— near-duplicate symbol pairs (explainshows what differs); diff-scopeshow— extract symbol body from source; Phase 13.3 truncation flags (signature_only/head/no_body/collapsed, mutually exclusive)outline— file structureusages— find all references to a symbol;filter_path/strict/whyimpact— delete-safety blast radius (verdict + per-channel evidence);depth/exclude_docsgrep— regex content searchpattern— AST pattern matching with metavar back-references; diff-scope;--whyimplementations— find types extending a base class/trait/interface (incl. generics); diff-scopesubtypes— transitive-down closure over extends/implements edges (direct children, grandchildren, …), depth-labelled; index-only (no live-walk fallback);depth/ diff-scopemodules— de-facto modules: clusters of symbols that call/reference each other (v9 index, computed on fullvex index); list clusters (label, size, cohesion, hubs) or passsymbolfor its cluster;limit/min_size/members/sort/ scope /workspace; empty result +empty_reasonon older indexes or--no-clusterscallers/callees— direct callgraph navigation (fast path via persistent index); diff-scopepaths— enumerate caller chains between two functionsreachable— transitive callers of a targettests_for— test functions that transitively cover a target (framework-labelled)diff— symbol-level diff between a git revision and the working treecheck— fast symbol existence checkbundle— unified multi-source bundle (mode: symbol | pr-impact | project), Phase 13 envelopeeval— ranking-evaluation harness (bench/min_ndcg), MCP defaultsjson: trueso agents get a structuredEvalReportcapabilities— machine-readable capability matrix (protocol_version,signals,bundle_modes,history_diff(v1.16.0),symbol_clusters, etc.)index/update— build/rebuild index; v1.16.0 addsgpu: bool/device: cpu|auto|cuda|directml|coremlargs (GPU-enabled builds only) so an agent can opt into GPU semantic embedding per-call without touching env or configstatus— index statistics (now includesgpu_support/default_device(v1.16.0))history— historical versions of a symbol across commits (MCP tool since v1.20.0, D5);depth/limit/since/until/author/kind/diff/exact_presence
Note:
vex historyandvex tests-forwere promoted to first-class MCP tools in v1.20.0 (D5); earlier docs that calledhistory"CLI-only" are stale. Both emit the same--format jsonenvelope as every other vex command.
Multi-repo (v1.22.0): eleven tools —
search,grep,check,usages,impact,callers,callees,reachable,modules,index,update— take aworkspace: booleanarg that fans the call across every.vex-workspace.tomlmember, returning the grouped{workspace, repos:[...]}payload understructuredContent.results. Pointproject_rootat or above the.vex-workspace.toml.find_symbolis excluded (usecheck/search);whyis ignored in workspace mode. Seedocs/MULTIREPO-PHASE8-mcp.md.
MCP ↔ CLI parity (v1.10): the schemas now mirror the CLI surface for every path-aware tool. Glob filters (include / exclude), substring filter, kind boost, context_path proximity hint, no_bm25, Phase 13.3 truncation, diff-scope, and no_stale_check are exposed everywhere the CLI accepts them — agents no longer need to drop to bash for "Rust files under crates/api/ since main"-style scoping.
The schemas follow a canonical vocabulary (query / symbol / symbols / path / pattern / filter / include / exclude); pre-v1.7 aliases (name, file, names, etc.) still work and emit _meta.deprecated_args: [...] in the JSON-RPC response. Malformed JSON-RPC input now returns the spec-compliant -32700 Parse error response (v1.9.2 fix) with a 512-codepoint echo of the offending line in the data field; broken-pipe / EOF on stdin cleanly shuts down the server instead of dropping in-flight tool calls. See docs/MCP-SCHEMA.md.
For other MCP-compatible clients (Cursor, Codex CLI, Windsurf, Cline, Continue.dev, Zed), see Other MCP Clients below — same vex-mcp binary, different config files.
Other MCP Clients
The same vex-mcp binary works with any MCP-compatible client. The binary install is identical to the Claude Code section above; only the per-client config file location and format differ.
One-line setup (v1.15.0+):
vex mcp install --agent cursor # or any of: claude-code, codex-cli, windsurf, cline, continue, zed
vex mcp install --agent all # fan out across every supported agent
vex mcp install --agent cursor --dry-run # preview the post-merge config without writingFor file-based agents, vex mcp install reads your existing agent config, merges a single vex server entry without disturbing siblings, and writes back atomically. For Claude Code it runs claude mcp add instead of editing a file. Idempotent — re-running on a matching entry is a no-op skip (--force overrides). vex mcp uninstall --agent <X> removes the entry; vex mcp list enumerates current entries per agent. The config files documented below for the other agents are exactly what vex mcp install writes — keep integrations/ handy for manual edits, agents the auto-installer doesn't know yet, or anything more exotic than the default shape.
Copy-pasteable snippets for the most common ones live under integrations/:
Agent | Snippet | Target file on disk |
Claude Code |
| registered via |
Cursor |
| |
Codex CLI (OpenAI) |
| |
Windsurf (Codeium) |
| |
Cline (CLI) |
| |
Continue.dev |
| |
Zed |
|
Per-agent caveats (auto-approve flags, timeout overrides, agent-mode requirements) are documented in integrations/README.md.
MCP Registry (from v1.27.2): vex is listed in the official MCP Registry as io.github.tenatarika/vex. Each release attaches one MCP Bundle per platform (vex-mcp-<target>.mcpb, macOS arm64 / Linux x86_64 / Windows x86_64) that holds both vex-mcp and vex; an MCPB-capable client asks for the project root once and needs nothing else on PATH. Bundle installs are updated by the client, not by vex self-update (which refuses to run inside a bundle).
Agent Recipes & Workflows
Once vex-mcp is wired into your agent, the next question is what to ask the agent so it picks the right tools in the right order. docs/COOKBOOK.md is a recipe collection for the common chains — code archaeology, cross-file refactor with usages --strict verification, PR-impact analysis via bundle(mode="pr-impact"), dead-code & duplicate cleanup, and multi-repo orchestration. Each recipe shows the tool sequence, the why of the ordering, and a phrase that reliably triggers the chain in agent prompts.
Documentation & Integration:
Full vex documentation and API reference: https://context7.com/tenatarika/vex (on Context7)
Agent skill reference:
.claude/skills/vex/SKILL.md(included in the project)
Shell Integration
# Shell completions (tab-completion for commands and flags)
vex completions bash > ~/.bash_completion.d/vex # Bash
vex completions zsh > ~/.zfunc/_vex # Zsh (add ~/.zfunc to fpath)
vex completions fish > ~/.config/fish/completions/vex.fish # Fish
# Aliases — add to .zshrc / .bashrc
alias vx="vex search"
alias vxu="vex usages"
alias vxi="vex index --path ."
alias vxs="vex index --path . --semantic"
alias vxw="vex watch"CLAUDE.md Integration
Add this to your project's CLAUDE.md to make Claude Code use vex instead of grep:
## Code Search
Before first use in a project, run `vex init` to generate `.vex.toml`, then `vex index` to build the index.
Set `auto_update = true` in `.vex.toml` so the index stays fresh automatically.
Use vex for code search instead of grep or manual file reading:
- `vex check "SymbolName"` — exact-name lookup: does it exist? (~4ms)
- `vex search "SymbolName"` — fuzzy symbol search: find definitions by name or meaning
- `vex search "description" --semantic` — search by meaning (requires --semantic index)
- `vex search "rare_term"` — BM25 channel finds rare terms in symbol bodies (auto-on when index has BM25 data)
- `vex show "SymbolName"` — extract symbol body (use INSTEAD of Read for specific symbols)
- `vex show "A" "B" "C"` — extract multiple symbols at once
- `vex usages "SymbolName"` — find all references
- `vex usages "SymbolName" --strict` — refactor-grade refs (binder-resolved, high precision)
- `vex impact "SymbolName"` — delete-safety blast-radius report (safe/unsafe/uncertain)
- `vex modules [SYMBOL]` — de-facto code clusters (symbol communities)
- `vex pattern 'class $NAME(BaseModel):' --lang python` — AST pattern matching with metavariables
- `vex pattern 'fn $N($$ARGS) -> Result<$T, $E> { $$$BODY }' --lang rust` — multi-line `$$$BODY` / `$$ARGS` capture
- `vex pattern 'struct $S && impl $S' --lang rust` — AND composition (back-ref `$S` must agree across both shapes)
- `vex pattern 'interface $N || class $N' --lang typescript` — OR composition (union, deduped by `(path, line)`)
- `vex pattern '<pat>' --lang <lang> --why` — emit ScanTrace on stderr (mode / candidate vs total / fallback reason)
- `vex outline path/to/file.py` — file structure overview
- `vex implementations "BaseService"` — find types extending a class/interface
- `vex subtypes "BaseService"` — transitive-down closure over extends/implements edges (direct children, grandchildren, …)
- `vex callers "function_name"` — find all callers (~4ms via persistent call graph)
- `vex callees "function_name"` — find all callees (~4ms via persistent call graph)
- `vex paths "from" "to"` — enumerate caller chains between two functions (multi-hop)
- `vex reachable "Target"` — transitive callers of a target (blast-radius analysis)
- `vex tests-for "SymbolName"` — test functions that cover a symbol (framework-labeled)
- `vex history "SymbolName"` — historical versions of a symbol across commits
- `vex similar "SymbolName"` — semantically close symbols (requires --semantic index)
- `vex duplicates --threshold 0.95` — near-duplicate symbol pairs
- `vex diff --base main` — symbol-level diff against a branch (added / removed / moved / body-changed)
- `vex bundle --mode symbol --symbol Foo` — single-call body + callers + callees + similar (replaces 4 round-trips)
- `vex bundle --mode pr-impact --base origin/main` — changed symbols + transitive callers + tests on the current branch
Many search-shaped commands support `--filter-path "path/"` (alias `--filter`) to narrow results to a directory (e.g. `search`, `show`, `usages`, `grep`, `similar`, `duplicates`). Most search-shaped commands also accept `--since <rev>` / `--since-branched` / `--changed-only` for diff-scoping.
### Rules
- **Always prefer `vex show` over `Read`** when you need a specific function or class
- **Always prefer `vex search` over `Grep`** when looking for symbol definitions
- **Use `vex grep` instead of `Grep`** for searching inside string literals, comments, or config values
- **Use `--format compact`** for token-efficient output in automated workflows
- **Use `--kind fn`** to boost results matching a specific symbol kind (fn, struct, trait, class, etc.)
- **Use `--context-path`** with the path of the file you are currently editing to boost nearby results
- **Run `vex update` after modifying source files** if `auto_update` is not enabled in `.vex.toml`
- **Use `vex pattern ... --why`** to debug match counts — the trace tells you whether the indexed prefilter ran or fell back to live-scan, and why
- **Indexed pattern prefilter requires a full `vex index`** — after `vex update` the section is partial and `vex pattern` automatically degrades to live-scan (reason `partial-section` in `--why`)
### Indexing
- `vex index` — full structural index + pattern skeleton section (v6)
- `vex index --semantic` — with embeddings (slower, enables semantic search)
- `vex update` — incremental update (only changed files)
- `vex index --no-pattern-index` — skip the v6 pattern skeleton section if you don't use `vex pattern` (sticky across `vex update`)
- `vex index --no-clusters` — skip computing symbol clusters (v9). `vex update` keeps the opt-out, but the next plain `vex index` computes clusters againTesting
Unit & Integration Tests
cargo nextest run --workspace # ~4,000 tests — unit, integration, property-based, adversarial
cargo test --doc # doctests
cargo clippy -- -D warnings # zero warnings policy(nextest ≥ 0.9.145 is recommended — older versions report spurious LEAKs on macOS; update with cargo nextest self update.)
Test coverage includes:
Per-language grammar regression (NEW):
tests/<lang>_query_test.rsfor all 19 supported languages — catches ABI mismatches and AST node renames when a tree-sitter grammar crate is upgradedBinary format: roundtrip, corrupted/truncated/wrong-version rejection, out-of-bounds access, string pool dedup, empty index
Adversarial format: 20 crafted index tests — overflow offsets, bad magic/version, alignment attacks, truncated records
Vectors: write/read roundtrip for 384-dim f32 embeddings
FST: refs FST roundtrip, prefix search, symbol FST exact/prefix/fuzzy search
Search: structural, fuzzy (Levenshtein), RRF fusion, reranking with kind/path/proximity boosts
Reranking stress: NaN/Infinity/zero scores, 10K results, edge context paths
Property-based (proptest): rerank preserves length, sorted output, no NaN/negative scores, fusion commutativity
Incremental update: unchanged reuse, deleted removal, file rename, symbol move between files, empty file
Concurrency: parallel index/update (lock serialization), concurrent readers, read during reindex
Multi-language: Rust, Python, Go, Kotlin, TypeScript, C++, cross-language same-name, wrong extension, 1K-symbol file, deep nesting, error recovery
Unicode: BOM, mixed CRLF, unicode identifiers, null bytes, empty/whitespace files
Path edges: spaces in paths, deep nesting (20 levels), symlinks, absolute vs relative, Windows backslashes
Callgraph: callers/callees for Rust, Python, Go, TypeScript, Java
Persistent call graph (v1.5): format v4 roundtrip, callers/callees FST lookup, dedup, same-name-across-files isolation, same-name-within-file disambiguation, incremental update preserves edges, fallback to live scan for v3
Similar/duplicates (v1.5): self-exclusion, threshold filtering, canonical pair dedup, body-length filter, empty-index handling
Pluggable embedder (v1.5): registry lookup, mismatch detection (incl. back-compat for pre-9.1 manifests), config + CLI priority, writer variable
vector_dimBM25 channel (v1.5): writer/reader roundtrip, pipeline emission, IDF discrimination, short-doc preference, 3-way RRF with Hybrid labeling, MatchType tagging, unicode tokens
Staleness: git HEAD comparison, dirty file detection, mtime fallback
Fuzz Testing
Fuzz tests exercise every parser that consumes untrusted input — the binary index format, sidecar files, the user-facing pattern grammar, and the JSON manifest — using cargo-fuzz (libFuzzer + AddressSanitizer):
# Install (once)
cargo install cargo-fuzz
# Generate seed corpus for every target
bash fuzz/generate_seeds.sh
# Run (requires nightly)
RUSTUP_TOOLCHAIN=nightly cargo fuzz run fuzz_index_reader -- -max_total_time=120
RUSTUP_TOOLCHAIN=nightly cargo fuzz run fuzz_refs_fst -- -max_total_time=60
RUSTUP_TOOLCHAIN=nightly cargo fuzz run fuzz_symbol_fst -- -max_total_time=60
RUSTUP_TOOLCHAIN=nightly cargo fuzz run fuzz_bloom_load -- -max_total_time=60
RUSTUP_TOOLCHAIN=nightly cargo fuzz run fuzz_pattern_parser -- -max_total_time=60
RUSTUP_TOOLCHAIN=nightly cargo fuzz run fuzz_manifest_load -- -max_total_time=60
RUSTUP_TOOLCHAIN=nightly cargo fuzz run fuzz_marker_load -- -max_total_time=60
RUSTUP_TOOLCHAIN=nightly cargo fuzz run fuzz_tokenize_document -- -max_total_time=60
RUSTUP_TOOLCHAIN=nightly cargo fuzz run fuzz_hash_index_load -- -max_total_time=60
RUSTUP_TOOLCHAIN=nightly cargo fuzz run fuzz_incremental_hnsw -- -max_total_time=60
RUSTUP_TOOLCHAIN=nightly cargo fuzz run fuzz_rename_chains_load -- -max_total_time=60
RUSTUP_TOOLCHAIN=nightly cargo fuzz run fuzz_state_load -- -max_total_time=60Eighteen fuzz targets cover the reader's unsafe paths plus every text /
sidecar parser that takes adversarial input:
Target | What it fuzzes | Surface |
| Arbitrary bytes as |
|
| Arbitrary FST + posting bytes |
|
| Arbitrary FST + posting bytes |
|
| Arbitrary |
|
| Arbitrary UTF-8 as a pattern string |
|
| Arbitrary JSON as |
|
| Arbitrary text as |
|
| Arbitrary UTF-8 as BM25 input |
|
| Arbitrary bytes as |
|
| Adversarial |
|
| Arbitrary bytes as |
|
| Arbitrary bytes as |
|
| Arbitrary FST + posting + edge bytes |
|
| Arbitrary bytes as Kotlin source | tree-sitter parse → symbol extraction → Kotlin |
| Arbitrary FST + posting + edge bytes |
|
| Arbitrary |
|
| Arbitrary bytes decoded as a ≤256-node graph | deterministic Leiden-CPM (run twice: identical output, every cluster connected) |
| Arbitrary bytes opened as an index file |
|
Most recent system-wide audit (Q4-A/B closure, 2026-06-17): ~76M
total executions across all 11 targets that existed at the time, 0 crashes / panics /
AddressSanitizer hits / leaks (61s per target, libFuzzer +
ASan + nightly). One latent FST panic was caught en route in a
dead-code path on adversarial ref_edge bytes and fixed before the
clean run — find_ref_edges_by_symbol and RefReader::find_by_prefix
now wrap the inner FST walk in catch_unwind so a corrupt sidecar
returns Err instead of taking down the process. Earlier baselines:
v1.14.1 (2026-06-05) ran 5.8M iterations across 9 targets clean;
v1.15.2 release-gate (2026-06-08) ran ~853k focused executions on
the four highest-signal targets clean.
Fuzzing has found and fixed eight real defects across the project life:
v1.x: out-of-bounds read on crafted
symbol_count, misaligned pointer dereference on oddsymbols_offset, unchecked section offsets exceeding file size (binary reader hardening).v1.12.0:
SymbolBloom::loadaccepted a sidecar withn_bits = 0k_num = 0whose consistency guard passed but later panicked insidebloomfilter::Bloom::checkonhash % 0. Fix: reject degenerate sizes during load.
v1.12.0:
SymbolBloom::loadacceptedk_numup to ~2.1B, which made everymay_containcall loop for 110+ seconds (DoS, not a panic). Fix: capk_num <= MAX_K_NUM = 64at load time.Phase 11.1.10 / Q4-A (2026-06-17): FST walk inside
find_ref_edges_by_symbolpanicked on a crafted refs FST, bypassing the productioncatch_unwind(the libfuzzer-sys panic hook fires before user code can intercept). Fix: wrap the inner walk in its owncatch_unwindand surfaceErr. Defense-in-depth hardening was also applied toManifest::load: a 128 MiB pre-read size cap blocks hostile JSON before serde can allocate multi-GB heap (Q4-B audit follow-up; defense-in-depth, threat model is user-owned files).v1.23.0: a 451-byte malformed Kotlin input drove tree-sitter's GLR error recovery into super-linear time and memory (334 s, >2 GB; DoS, not a crash; found by
fuzz_kotlin_binder). Fix: every production parse goes throughparser_pool::parse_text, which caps progress-callback invocations (scaled by input size) for all languages.v1.23.0:
tree_sitter::Node::utf8_text()panicked on malformed input where tree-sitter emitted a node past EOF (found byfuzz_kotlin_binder). Fix: the bounds-checkedNodeTextExt(node_text/node_text_opt) replaces rawutf8_textin the extractor, binders and pattern prefilter.
The v1.13.0 / v1.14.1 additions found no defects in fresh code — the
review-driven MAX_COUNT guards on hash_index::save / load were
added as defence-in-depth before the fuzzer ran (rust-reviewer +
code-reviewer flagged the truncating as u32 cast on save), and the
sustained 3M / 5.8M iteration runs confirmed they hold.
Architecture
CLI (clap) → Pipeline (rayon, 500-file chunks) → Tree-sitter
↓
Binary format v9 (mmap, zero-copy)
↓
┌──────────────────┬──────────────┬──────────────┬──────────────┬─────────────┬──────────────┐
↓ ↓ ↓ ↓ ↓ ↓ ↓
Symbol FST Refs FST BM25 doc HNSW vectors Call graph Hierarchy Clusters
(structural) (cross-file refs) (body tokens) (semantic) (callers FST / (v8 edges) (v9 Leiden-
callees CSR) CPM)
↓
Embedder trait → fastembed / MiniLM-L6 (default)
Per-project sidecars (in <index_dir>/):
· index.vex — primary index (format v9)
· manifest.json — metadata: embedder, sections, version, staleness tracking
· index.bloom — symbol-name bloom filter (`vex check` skips FST lookups for definitely-missing names)
· index.trigram — per-file trigram bloom so `vex grep` skips non-matching files (v1.24.1)
· index.hnsw — semantic vectors (HNSW graph)
· index.bodytokens — per-symbol terms for BM25 + semantic context (B1.2)
· index.git_history — historical symbol presence (Phase 14.8, FST + git-walk fallback)
· index.rename_chains — MinHash+LSH rename tracking across commits (Phase 14.10)
· index.state — incremental state: imported_by reverse map + writer-provenance sentinels (audit C1)
Shared cross-project (in user cache root, e.g. ~/Library/Caches/vex/blobs/):
· {sha}.bin shards — content-addressed parse cache, keyed by git blob SHA (Phase 14.7)
Search pipeline:
search → Symbol FST + BM25 + HNSW → N-way RRF fusion → Hybrid tag on cross-channel hits
usages → Refs FST + posting lists → zero-copy, --strict adds type-aware filter (v1.14.1)
history → walks git tree, follows rename chains via Jaccard + greedy 1:1 (Phase 14.10)
similar → HNSW nearest neighbors (hash-keyed for content-stable IDs)
show → tree-sitter node boundaries → symbol body extraction
pattern → AST-aware structural matcher with skeleton indexNo SQLite — custom binary format v6, zero-copy mmap reads; readers accept v3+ for backwards compatibility
Symbol FST — persistent inverted index, O(query_len) lookup
Refs FST + ref_edges — symbol references as FST + cross-file edges resolved at write time (Pass-2 in
store::writer); enables refactor-gradeusages --strictPersistent call graph —
CallEdgerecords + a name-keyed callers FST + a dense callees CSR index (v9+; FST on older indexes), built at index time, ~4ms lookup vs seconds of live tree-sitter scanBM25 channel — Okapi BM25 over
body_tokens, auto-on when section presentHNSW — approximate nearest neighbor via usearch, O(log N) semantic search; hash-keyed entries for content-stable IDs across re-indexing
Pluggable embedder —
Embeddertrait + registry, identity recorded in manifest with mismatch detection at searchHistory index — symbol presence per commit + MinHash-based rename chain tracking (closes LIMITATIONS §4c #2 for 1:1 renames)
Parallel parsing — rayon with 500-file chunks; blob-SHA parse cache (shared across projects in the user cache root) skips re-parse of unchanged files across re-indexes
Incremental updates — content hashing via xxh3;
vex updatere-parses only changed files (unchanged symbols + call edges reconstructed from existing index)Watch mode —
notifycrate with 500ms debouncingN-way RRF fusion —
fuse_manymerges structural + BM25 + semantic ranked lists, marks cross-channel hits asHybridRanking eval harness —
vex evalover a bundled golden set; CI regression gate on mean nDCG@10 + per-query-type floors + per-channel attribution (Phase 13.12 / 13.12.1)
License
MIT
Available Tools
28 toolsbundleA
Multi-source bundle — replaces 4 round-trips (show → callers → callees → similar) with 1. Three modes: symbol (body + callers + callees + similar for a named symbol; ~10ms), pr-impact (changed symbols + transitive callers + tests for a git base ref; ~50ms), project (top-N symbols by reverse call-graph indegree; ~5ms). Prefer over chaining find_symbol/show/callers/callees when you need cross-section context on one symbol or a PR. Mode-specific args are validated server-side; only mode is universally required. Response shape is uniform — { protocol_version, capabilities, _meta, results: { mode, items[], mode_hints } }. Each items[i] carries 13.11 signals plus a role discriminator (body | caller | callee | similar | changed | transitive_caller | test | top). Scope filters (include / exclude / exclude_tests) apply only in pr-impact mode (changed files plus the caller and test rows); symbol and project modes ignore them.
| Name | Required | Description | Default |
|---|---|---|---|
| base | No | (mode: pr-impact) Git base revision to diff against (e.g. `origin/main`, `HEAD~3`, a SHA) | |
| mode | Yes | Bundle assembly mode | |
| depth | No | (mode: pr-impact) Transitive callers walk depth | |
| top_n | No | (mode: project) Max number of top-ranked symbols | |
| symbol | No | (mode: symbol) Symbol name to resolve via the symbol FST | |
| exclude | No | Blacklist results by path glob; wins over include (repeatable) | |
| include | No | Whitelist results by path glob (repeatable) | |
| path_glob | No | (mode: project) Single path glob filter applied to ranked symbols (e.g. `src/**`); separate from the universal `include`/`exclude` arrays | |
| tests_max | No | (mode: pr-impact) Max test-classified items | |
| auto_update | No | Auto-update the index if stale, or bootstrap if missing, before running (default: true) | |
| callees_max | No | (mode: symbol) Max direct callees | |
| callers_max | No | (mode: symbol) Max direct callers | |
| similar_max | No | (mode: symbol) Max semantic-similar matches; gated on `vex index --semantic` | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations the description carries the full burden, and it delivers unusually rich context: per-mode latency, server-side arg validation, the uniform `{ protocol_version, capabilities, _meta, results }` envelope, the `role` discriminator values, and that scope filters only apply in pr-impact mode. The main gap is that it doesn't flag the state-changing side of auto_update/async_update (index refresh) or any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and front-loaded: value proposition first, then modes, then response contract, then the filter-scope caveat. It is long, but every sentence carries information an agent needs; only the timing figures are marginally expendable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-parameter, three-mode tool with no output schema, the description compensates by documenting the response envelope, item structure, and role vocabulary. Combined with 100% schema coverage, an agent has what it needs to select a mode and call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (baseline 3), and the description adds cross-parameter semantics the schema can't express well: which mode each argument belongs to and the fact that include/exclude/exclude_tests are ignored outside pr-impact. It doesn't add format examples beyond what the schema already supplies inline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening states a concrete verb+resource ('Multi-source bundle') and immediately frames it against the exact sibling tools it replaces (show → callers → callees → similar). The three modes are named and each is scoped to a distinct use case, so an agent can distinguish it from find_symbol, callers, callees, and similar without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says 'Prefer over chaining find_symbol/show/callers/callees when you need cross-section context on one symbol or a PR,' giving both the alternative and the selecting condition. Mode-specific guidance (symbol = one symbol, pr-impact = a git base ref, project = ranked symbols) further routes the agent to the correct mode.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calleesA
Direct callees of a function via the persistent call-graph FST (~4ms when indexed; falls back to live-scan). Prefer over Read+manual scanning when you want to know what a function calls without reading the whole body — callees gives the resolved outgoing edges as records. Phase 14.2 + 14.2.2 + 14.2.1: Python/Java decorators, Kotlin annotations, C# method/constructor attributes, TypeScript method decorators, and Rust outer attributes on fns/methods are surfaced as callees of the decorated function (decorator factories like @lru_cache(maxsize=128), @Inject, @Get("/x"), or #[tokio::test] appear as the path-rightmost identifier lru_cache / Inject / Get / test alongside regular body calls). Rust #[derive(...)] is intentionally filtered. Supports diff scoping: since / since_branched / changed_only (mutually exclusive) to restrict callees to recently-touched code.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | DEPRECATED — use `symbol`. Pre-v1.7 alias, still accepted; emits a deprecated_args notice in _meta. | |
| limit | No | Max results | |
| since | No | Restrict results to files changed between `<rev>..HEAD` (accepts anything `git diff` understands: `main`, `HEAD~3`, `origin/main`, SHA). Mutually exclusive with `since_branched` and `changed_only`. | |
| symbol | Yes | Exact function name — canonical key (v1.7+). | |
| exclude | No | Blacklist results by path glob; wins over include (repeatable) | |
| include | No | Whitelist results by path glob, gitignore syntax (repeatable) | |
| workspace | No | Multi-repo: fan out across every repo declared in the nearest `.vex-workspace.toml` (set `project_root` at or above it — the manifest is found by walking up). Results become an object `{workspace, repos:[...]}` grouped by repo, NOT the flat per-tool array — branch on shape. `why` is ignored in workspace mode (single-repo only). | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running — enables the call-graph fast path (default: true) | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| changed_only | No | Restrict results to working-tree changes (staged + unstaged + untracked). Mutually exclusive with `since` and `since_branched`. | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true. | |
| since_branched | No | Restrict results to files changed since this branch diverged from `origin/main` (or `main`/`master`). Mutually exclusive with `since` and `changed_only`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so the description carries the full burden. It discloses latency (~4ms indexed, fallback to live-scan), decorator/annotation surfacing behavior for 6 languages (with concrete examples), that Rust #[derive] is intentionally filtered, and diff scoping semantics. Solid behavioral disclosure, though it doesn't cover auth, rate limits, or error modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose well, but is quite long and includes a verbose Phase-number changelog block (14.2 + 14.2.2 + 14.2.1) that reads like release notes rather than tool-selection guidance. The decorator detail is useful but the enumeration could be tightened significantly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-param, no-output-schema tool, the description covers fallback behavior, result format hints ('resolved outgoing edges as records'), workspace shape-changing behavior is in schema, and diff scoping. Complete enough to call correctly; could say more about output record shape since there's no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all 14 parameters thoroughly. The description reiterates 'since/since_branched/changed_only (mutually exclusive)' which is also in the schema. Baseline 3 is correct when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb/resource: 'Direct callees of a function via the persistent call-graph FST'. This clearly distinguishes it from siblings like 'callers' (inverse direction) and 'Read+manual scanning'. The parenthetical about fallback behavior adds specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: 'Prefer over Read+manual scanning when you want to know what a function calls without reading the whole body'. This gives an explicit alternative and the condition selecting this tool. No explicit when-not-to-use, but strong positive guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
callersA
Direct callers of a function via the persistent call-graph FST (~4ms when indexed; falls back to live-scan). Prefer over grep for who calls Foo? — grep on the function name hits doc comments and string literals; the call-graph edges are resolved at parse time. Phase 14.2 + 14.2.2 + 14.2.1: Python/Java function/method decorators, Kotlin annotations / C# method+constructor attributes, and TypeScript method decorators / Rust outer attributes on fns/methods emit forward edges, so callers GetMapping lists every Spring handler, callers get lists every FastAPI route, callers HttpGet every ASP.NET action, callers JvmStatic every Kotlin function annotated @JvmStatic, callers Get every Nest.js @Get(...), callers test every Rust #[tokio::test] (the rightmost identifier of the decorator/attribute path becomes the callee; arguments are ignored — #[serde(rename = "x")] → serde, not rename). Rust #[derive(...)] is filtered (compile-time codegen, not call edges). Note the rightmost-identifier convention means callers get mixes decorator handlers with any regular .get() call — narrow with include/exclude if needed. Pair with paths for multi-hop chains. Supports diff scoping: since / since_branched / changed_only (mutually exclusive) to restrict callers to recently-touched code.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | DEPRECATED — use `symbol`. Pre-v1.7 alias, still accepted; emits a deprecated_args notice in _meta. | |
| limit | No | Max results | |
| since | No | Restrict results to files changed between `<rev>..HEAD` (accepts anything `git diff` understands: `main`, `HEAD~3`, `origin/main`, SHA). Mutually exclusive with `since_branched` and `changed_only`. | |
| symbol | Yes | Exact function name — canonical key (v1.7+). | |
| exclude | No | Blacklist results by path glob; wins over include (repeatable) | |
| include | No | Whitelist results by path glob, gitignore syntax (repeatable) | |
| workspace | No | Multi-repo: fan out across every repo declared in the nearest `.vex-workspace.toml` (set `project_root` at or above it — the manifest is found by walking up). Results become an object `{workspace, repos:[...]}` grouped by repo, NOT the flat per-tool array — branch on shape. `why` is ignored in workspace mode (single-repo only). | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running — enables the call-graph fast path (default: true) | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| changed_only | No | Restrict results to working-tree changes (staged + unstaged + untracked). Mutually exclusive with `since` and `since_branched`. | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true. | |
| since_branched | No | Restrict results to files changed since this branch diverged from `origin/main` (or `main`/`master`). Mutually exclusive with `since` and `changed_only`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the burden and does well: it discloses the fast-path timing (~4ms indexed), live-scan fallback, parse-time edge resolution, and the decorator rightmost-identifier convention including the `serde` vs `rename` edge case and filtered `#[derive(...)]`. It omits any mention of index side effects (auto_update bootstrap writes an index) and auth/permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and the grep comparison are front-loaded, but the middle devolves into a long run-on sentence enumerating five near-identical framework examples that restate the same rightmost-identifier rule. Dense and useful, but the example list is repetitive rather than earning each clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-param read tool with no annotations and no output schema, coverage is strong: diff-scoping exclusivity, decorator semantics, and workspace shape are all addressed. The remaining gap is return-format behavior beyond the workspace note, which the missing output schema leaves to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% so the baseline is 3, but the description adds real semantic value the schema lacks: the rightmost-identifier convention that determines what `symbol` matches, the decorator/attribute edge sources, and how that convention makes `callers get` mix decorator handlers with `.get()` calls.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource: 'Direct callers of a function via the persistent call-graph FST'. It explicitly distinguishes itself from grep and names siblings (paths, tests_for) so an agent can route correctly without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'Prefer over grep for `who calls Foo?`' with the reason (grep hits doc comments/string literals), plus guidance to narrow with include/exclude and to pair with paths for multi-hop chains. When-to-use and alternatives are both stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capabilitiesA
Return vex protocol version + capability matrix for client capability negotiation.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It does disclose the return payload (protocol version + capability matrix), which is the key behavioral fact and implies a side-effect-free read. It does not mention auth requirements, caching, or whether negotiation is idempotent, leaving some gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the verb first and zero wasted words. Ideal structure for a trivial no-arg endpoint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with no output schema, the description names what is returned but not the shape of the capability matrix. Adequate to call the tool, but an agent cannot anticipate the response structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which is the baseline-4 case. There is nothing to document and the description correctly implies a no-argument call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Return') and resource ('vex protocol version + capability matrix'), making the tool's function unambiguous. It does not explicitly name a sibling to differentiate from, but the sibling set is entirely code-analysis tools, so confusion is unlikely.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'for client capability negotiation' implies the context of use, which is a reasonable cue. However, there is no explicit when/when-not guidance or named alternative, so the agent must infer the trigger condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
checkA
Batch existence probe: confirm whether one or more symbol names exist in the index without paying for body extraction or ranked search (~4ms total). Use before show / usages / callers when working from an unverified list — skip the symbols that don't exist instead of letting downstream tools error.
| Name | Required | Description | Default |
|---|---|---|---|
| names | No | DEPRECATED — use `symbols`. Pre-v1.7 alias, still accepted; emits a deprecated_args notice in _meta. | |
| symbols | Yes | Exact symbol names to probe — canonical key (v1.7+). | |
| workspace | No | Multi-repo: fan out across every repo declared in the nearest `.vex-workspace.toml` (set `project_root` at or above it — the manifest is found by walking up). Results become an object `{workspace, repos:[...]}` grouped by repo, NOT the flat per-tool array — branch on shape. `why` is ignored in workspace mode (single-repo only). | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running (default: true) | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does well: it discloses cost (~4ms), that no body extraction or ranking happens, and that it prevents downstream errors. It omits that auto_update (default true) can bootstrap or mutate the index as a side effect, which is a meaningful behavioral gap for a tool framed as a cheap probe.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, with the core capability front-loaded before the routing guidance. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with no annotations and no output schema, the description covers purpose, cost, and routing well; workspace-mode result shape and index-refresh behavior live in the schema. The one unaddressed item is the index-mutation side effect implied by auto_update, which an agent might want called out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (including the deprecated names alias, workspace shape branching, and stale-check flags) is already documented in the schema. The description adds no parameter-level detail beyond what is structured, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (confirm existence) and resource (symbol names in the index), and explicitly contrasts itself with ranked search and body extraction. An agent can distinguish it from show/usages/callers/search without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('before show / usages / callers when working from an unverified list') and the reason ('skip the symbols that don't exist instead of letting downstream tools error'). Names the alternatives and the condition that selects this one.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
diffA
Symbol-level diff between a git revision and the working tree: lists added / removed / moved / body-changed symbols on the touched files. Prefer over git diff + manual scanning for PR review — git diff returns line hunks while this returns structured symbol records, so an agent can iterate over changed-functions directly instead of parsing unified-diff text.
| Name | Required | Description | Default |
|---|---|---|---|
| base | Yes | Git revision to compare against (e.g. main, HEAD~3, origin/main). Working tree is the new side. | |
| limit | No | Max changes to return | |
| exclude | No | Blacklist changes by path glob; wins over include (repeatable) | |
| include | No | Whitelist changes by path glob, gitignore syntax (repeatable) | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the output shape (structured symbol records, iterable directly) and the change taxonomy, which goes beyond the schema. It does not discuss ordering, whether results are paged/truncated relative to `limit`, or read-only guarantees, which are minor omissions for a diff tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both front-loaded: the first defines the operation and its output, the second routes the agent away from `git diff`. No filler or restated boilerplate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description correctly steps in to describe the return value as symbol records grouped by change type. It is sufficient to call the tool correctly; only minor gaps remain (ordering, how `limit` truncation is signaled).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and every parameter (base, limit, exclude/include precedence, project_root, exclude_tests) is fully documented in the schema itself. The description adds no parameter-level detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('symbol-level diff between a git revision and the working tree') plus the exact output categories (added/removed/moved/body-changed symbols). It distinguishes itself from the closest non-MCP alternative, `git diff`, and from line-hunk tools generally.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative ('git diff' + manual scanning) and the scenario that selects this tool ('PR review'), and explains why: structured symbol records vs. unified-diff text. Nothing is left to inference about when to reach for it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
duplicatesA
Repo-wide near-duplicate scan: pairs of symbols whose embeddings exceed threshold. Use for refactor planning (where else does this logic live?) and dedup. Prefer over manual similar-walks — duplicates evaluates all pairs once with min_body_lines filtering out trivial bodies. Requires vex index --semantic. Supports diff scoping: since (rev), since_branched, changed_only (mutually exclusive) and no_stale_check.
| Name | Required | Description | Default |
|---|---|---|---|
| why | No | Surface a JSON trace under `_meta.why`: applied threshold + min_body_lines, pairs before/after path filter, filter snapshot. | |
| limit | No | Max pairs to return | |
| since | No | Restrict pairs to files changed between `<rev>..HEAD`. Mutually exclusive with `since_branched` and `changed_only`. | |
| exclude | No | Blacklist pairs by path glob — a pair is dropped when either side matches (repeatable) | |
| explain | No | Include reasoning per pair: identifier-set Jaccard overlap + truncated unified diff between the two bodies | |
| include | No | Whitelist pairs by path glob — a pair is kept when at least one side matches (repeatable) | |
| threshold | No | Minimum cosine similarity (0.0..1.0); 0.9 keeps only very close pairs | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running (default: true) | |
| filter_path | No | Substring path filter — keep pairs where at least one symbol's path contains this substring. Legacy alias: `filter`. | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| changed_only | No | Restrict pairs to working-tree changes (staged + unstaged + untracked). Mutually exclusive with `since` and `since_branched`. | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| min_body_lines | No | Skip symbols with body shorter than this many lines (filters trivial wrappers) | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true. | |
| since_branched | No | Restrict pairs to files changed since this branch diverged from `origin/main` (or `main`/`master`). Mutually exclusive with `since` and `changed_only`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden well: it discloses the hard prerequisite ('Requires `vex index --semantic`'), the filtering behavior of min_body_lines, and the mutual exclusivity of the three diff-scoping flags. It stops short of describing the return shape (pair list) or how many pairs come back by default.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense, front-loaded sentences: purpose, usage, then prerequisites and scoping flags. Every clause carries information, though the final scoping sentence packs several flags together and reads a bit like a spec dump.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter, zero-required, annotation-free tool with no output schema, the description covers the prerequisite, the filtering model, and the diff-scoping exclusivity — the things an agent most needs. The missing piece is what a result actually looks like (pair structure/ordering), which nothing else supplies.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter semantics the schema cannot express — the mutual exclusivity of `since`/`since_branched`/`changed_only` and the fact that `no_stale_check` is redundant under `auto_update`. It restates threshold/min_body_lines meaning rather than extending it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource+scope: 'Repo-wide near-duplicate scan: pairs of symbols whose embeddings exceed threshold.' That clearly separates it from symbol-lookup siblings, though it never names the closest siblings (find_similar, similar) and only alludes to them as 'manual similar-walks.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete use cases ('refactor planning', 'dedup') and a comparison condition ('Prefer over manual similar-walks — duplicates evaluates all pairs once with min_body_lines filtering'). No explicit when-not guidance and no direct routing to find_similar/similar for the single-symbol case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
evalA
Run the ranking-quality harness against a golden query set and return nDCG@10 / recall@10 / MRR per query and aggregated. Indexless in the sense that it never builds — consumes whatever index already lives at the project root (run index first if missing). Intended as a CI regression guard. MCP defaults to json: true so agents receive structured EvalReport JSON instead of the human-readable summary the CLI emits.
| Name | Required | Description | Default |
|---|---|---|---|
| json | No | Emit the EvalReport as JSON to stdout. Default `true` in MCP context (agents want structured output) — note the CLI default is `false`. Set explicitly to `false` to fall back to the human-readable summary. | |
| bench | No | Path to the golden-set TOML. Defaults to the bundled `benches/ranking_golden/queries.toml` on the CLI side; pass this when running against a fixture. | |
| min_ndcg | No | Fail with non-zero exit if mean nDCG@10 drops below this floor. Default 0.0 (always succeed). CI pins a recorded floor. | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that the tool is indexless (never builds), consumes only the existing index, and that min_ndcg drives a non-zero exit — i.e. it can fail. It omits behavioral traits like runtime cost or idempotency, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, all substantive: action, indexless precondition, CI intent, and the MCP json default lead in order of importance. Slightly dense with parentheticals, but nothing is pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-param, no-output-schema tool with no annotations, the description covers action, deliverables, dependency ordering, failure semantics, and a default divergence between MCP and CLI. There is no output schema, but the return shape (per-query and aggregated metrics as EvalReport JSON) is named.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters, including the MCP-vs-CLI default for `json`. The description's parameter remarks largely restate schema content, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Run the ranking-quality harness against a golden query set") and enumerates the exact outputs (nDCG@10, recall@10, MRR). This clearly separates it from sibling tools like `index`, `tests_for`, or `check`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit context ("Intended as a CI regression guard") and a prerequisite ("run `index` first if missing"). It lacks an explicit when-not-to-use or a direct comparison against a sibling such as `check`, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_similarA
Semantic-only search by natural-language description (e.g. 'payment processing' → ChargeUseCase, BillingService). Uses the HNSW vector index built by vex index --semantic (~7-15ms). Prefer over search when you do not know any concrete identifier and want concept-level matching; prefer search when you have a partial name (search fuses semantic + lexical channels for better recall on identifier-shaped queries).
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Natural-language description of the concept (not an identifier; use find_symbol for those). | |
| exclude | No | Blacklist results by path glob; wins over include (repeatable) | |
| include | No | Whitelist results by path glob, gitignore syntax (repeatable) | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running (default: true) | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden; it discloses the underlying HNSW vector index, the ~7-15ms latency, and that the index is built by `vex index --semantic`. It does not describe return shape or result ranking, and staleness/update behavior is left largely to the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core purpose, then mechanism, then routing guidance. Dense but every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with no annotations and no output schema, the description covers purpose, routing, and index/latency behavior well. It omits return-value expectations and ranking behavior, but the routing and index context make it largely callable as-is.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all eight parameters are already documented, including the 'not an identifier' caveat on `query`. The description adds no parameter-level detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Semantic-only search by natural-language description') with a concrete example mapping. It explicitly distinguishes itself from `search` and `find_symbol`, so an agent can identify the tool without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit decision rule: prefer this over search when no concrete identifier is known, prefer search when a partial name exists, and use find_symbol for identifiers. Both the when and the when-not with named alternatives are present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_symbolA
Resolve a symbol by exact name (with prefix fallback) against the FST inverted index (~4ms). Prefer over search when the symbol name is known and you want exactly that record back, not a fused-rank list. Prefer over grep for git grep 'class Foo'-style definition lookup — grep scans every byte; this is a constant-time index probe.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | DEPRECATED — use `symbol`. Pre-v1.7 alias, still accepted; emits a deprecated_args notice in _meta. | |
| symbol | Yes | Exact symbol name (function/class/struct/etc.) — canonical key (v1.7+). Use search for partial or fuzzy names. | |
| exclude | No | Blacklist results by path glob; wins over include (repeatable) | |
| include | No | Whitelist results by path glob, gitignore syntax (repeatable) | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running (default: true) | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose real behavior: prefix fallback on exact match, index-based resolution, ~4ms latency, and constant-time probing vs. grep's byte scan. It omits that it is a read-only operation and any mention of index staleness/bootstrapping, but the core behavioral profile is well communicated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core verb+resource and mechanism, then routing guidance. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no output schema, the description covers the primary use case and alternatives well; the remaining parameters (include/exclude globs, auto_update, async_update, no_stale_check) are fully covered by the 100% schema. Return values needn't be explained since none are promised.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 9 parameters in detail. The description only reinforces the canonical `symbol` semantics ('exact name with prefix fallback') beyond that, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (resolve) and resource (a symbol by exact name) plus scope and mechanism (FST inverted index, ~4ms). It explicitly distinguishes itself from two siblings, search and grep, so an agent can choose without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-prefer conditions: over search 'when the symbol name is known and you want exactly that record back, not a fused-rank list', and over grep for definition-style lookup. It names the alternatives and the conditions that select them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
grepA
Regex content search across files (ripgrep-equivalent, no index needed). Use this for searching inside string literals, comments, config values, or any non-symbol text. Prefer search / find_symbol / usages for identifier lookups — those are index-backed (~4ms) while grep is a full-scan and returns raw line matches without symbol context.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | Force-read every file, bypassing the binary-file skip (extension denylist + NUL/high-control content sniff). Escape hatch for a legitimately-textual file that got misclassified as binary; a genuinely invalid-UTF-8 file is still skipped. CLI equivalent: `-a`/`--text` (ripgrep parity). | |
| limit | No | Max results | |
| exclude | No | Blacklist results by path glob; wins over include (repeatable) | |
| include | No | Whitelist results by path glob, gitignore syntax (repeatable) | |
| pattern | Yes | Regex pattern (Rust regex syntax) to match against file contents. | |
| workspace | No | Multi-repo: fan out across every repo declared in the nearest `.vex-workspace.toml` (set `project_root` at or above it — the manifest is found by walking up). Results become an object `{workspace, repos:[...]}` grouped by repo, NOT the flat per-tool array — branch on shape. `why` is ignored in workspace mode (single-repo only). | |
| filter_path | No | Substring path filter applied to result paths (single substring; use include/exclude for glob patterns). Legacy alias: `filter`. | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it delivers: no index required, full-scan cost, and that results are raw line matches without symbol context. It does not cover return shape or result-limit behavior beyond the schema, but the core behavioral profile is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with what the tool is, then usage, then the alternative-routing rationale. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, no-output-schema tool, the description conveys the search model and its relationship to index-backed siblings, which is what an agent needs to choose correctly. It leaves return-shape details implicit, but the schema covers parameter nuance thoroughly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 9 parameters (including nuanced ones like text, exclude_tests, workspace) are documented in the schema itself. The description adds no parameter-level meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (file contents) and adds the distinguishing mechanism (regex, ripgrep-equivalent, no index needed). An agent can tell it apart from index-backed siblings like search and find_symbol without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the when (string literals, comments, config values, non-symbol text) and the when-not, routing identifier lookups to search/find_symbol/usages. It even justifies the routing with the performance contrast (index-backed ~4ms vs full-scan).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
historyA
Every historical version of a symbol reachable from a chosen tip. With vex index --history previously run, queries hit a persistent FST sidecar (~ms); without it, shells out to git log (~seconds). Indexed mode also finds symbols whose name has been DELETED from HEAD — the walker can't. Use this to inspect how a function's body / signature changed over time, find when a bug was introduced, or recover a deleted symbol's last definition. NOTE: omitting limit returns the full history (walker mode is unbounded by default — set limit to cap latency on long-lived repos). exact_presence: true adds seconds-scale latency per file — only pass when you specifically need the exact commit set, not the convex-hull span. v1.20.0 (D5) surface — the CLI subcommand has existed since v1.15.0 but was MCP-invisible.
| Name | Required | Description | Default |
|---|---|---|---|
| diff | No | Render unified diffs between consecutive historical versions of the same `(symbol, kind)` pair instead of repeating the full body for each entry. Cuts output noise on deep histories. Mutually exclusive with `exact_presence`. | |
| kind | No | Keep only entries whose symbol kind matches exactly (lowercase: `function` / `struct` / `impl` / …). | |
| name | No | DEPRECATED alias for `symbol`; still accepted, emits a deprecated_args notice in _meta. | |
| depth | No | Max commits to walk per file (walker mode). Unbounded by default; bump down on long-lived repos to keep latency in check. | |
| limit | No | Cap the total result set. Omit for unbounded (walker mode) — set explicitly on long-lived repos to keep latency in check. The walker stops as soon as the limit is reached. | |
| since | No | Keep only entries whose commit date is `>= YYYY-MM-DD` (inclusive). | |
| until | No | Keep only entries whose commit date is `<= YYYY-MM-DD` (inclusive). | |
| author | No | Keep only entries whose commit author contains this substring (case-insensitive). Walker-only — the indexed path rejects this with an error pointing at `no_index: true`. | |
| branch | No | Restrict the walk to this revision (`refs/heads/foo`, `origin/main`, a SHA). Defaults to `HEAD`. | |
| symbol | Yes | Symbol name to walk through history. Matched whole-word via `git grep --word-regexp`, then filtered post-parse to exact `name == query`. | |
| no_index | No | Force the v1.16 query-time walker even when a `git_history` section is present. Default (`HistoryMode::Auto`) picks the indexed path when available and falls back to the walker otherwise. Use for regression-checking the walker against the indexed path. | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exact_presence | No | For each entry, list the exact set of commits where its blob lived in the file. Defeats the convex-hull span representation (LIMITATIONS §4c #4). Adds latency. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it distinguishes indexed FST (~ms) vs walker (`git log`, ~seconds) modes, discloses that indexed mode finds deleted symbols while the walker cannot, warns that omitting `limit` is unbounded, notes `exact_presence` latency, and explains `no_index` for regression checking. Missing only auth/permission context, which is minor here.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with scope and then performance/mode tradeoffs; most sentences earn their place. Slight redundancy – the unbounded-latency warning appears both in the description prose and in the `limit`/`depth` schema descriptions – keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, with 13 parameters, so the description must carry a lot, and it does: modes, latency, correctness differences, and history/version provenance. It stops short of describing the returned entry shape (versions, spans, convex-hull representation), which would help given the absent output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real value beyond the schema: the default-unbounded behavior of `limit`, the per-file latency of `exact_presence`, and the cross-parameter note that `diff` and `exact_presence` are mutually exclusive. It reinforces rather than repeats structured data.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource scope: 'Every historical version of a symbol reachable from a chosen tip.' It is clearly distinguishable from siblings like `find_symbol`, `diff`, or `callers`, and the extra sentence about deleted symbols sharpens exactly what makes this tool unique. An agent can route to it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives three concrete use cases (inspect body/signature change over time, find when a bug was introduced, recover a deleted symbol's last definition) and conditions for latency-sensitive flags. It does not name a sibling alternative to prefer, e.g. when `diff` or `find_symbol` would be the better call, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
impactA
Delete-safety blast-radius report. Composes four independent reference channels — strict refs (binder-resolved v5 edges), the legacy FST refs, grep \b<Name>\b against the project, and direct call-graph callers — into a single verdict (safe / unsafe / uncertain). Use this BEFORE proposing to delete or rename a symbol; one call collapses what CLAUDE.md previously documented as a manual dance across usages → grep → callers. Verdict rule: unsafe if strict_refs > 0 OR call_graph_callers > 0 (binder/graph confirmed real usage); uncertain if only text channels (FST / grep) hit (likely string-dispatch / decorator / comment mentions); safe only when every channel reports zero hits. results shape: { symbol, verdict, verdict_explanation, channels: { strict_refs, fst_refs, grep_word_boundary, call_graph_callers } } where each channel block has { available, count, sample[], truncated }.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | DEPRECATED — use `symbol`. Pre-v1.7 alias, still accepted; emits a deprecated_args notice in _meta. | |
| depth | No | (v1.21.0) BFS hop budget for transitive callers. `1` (default) reports direct callers only via `call_graph_callers`; `>= 2` enables the `transitive_callers` channel, walking the call graph backward up to N hops. Silently clamped to `[1, 16]`. Use to see the full upstream blast radius (`outer -> middle -> leaf` chain surfaces `outer` at depth=2). | |
| symbol | Yes | Exact symbol name to assess — canonical key. | |
| exclude | No | Blacklist results by path glob; wins over include (repeatable). | |
| include | No | Whitelist results by path glob, gitignore syntax (repeatable). Applied to every channel — useful for scoping to e.g. `src/**` when assessing a library symbol. | |
| workspace | No | Multi-repo: fan out across every repo declared in the nearest `.vex-workspace.toml` (set `project_root` at or above it — the manifest is found by walking up). Results become an object `{workspace, repos:[...]}` grouped by repo, NOT the flat per-tool array — branch on shape. `why` is ignored in workspace mode (single-repo only). | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running (default: true) | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| exclude_docs | No | (v1.20.1, D4 parity) Opt-in: drop text-channel hits in prose-format files (`*.md`/`*.markdown`/`*.txt`/`*.rst`/`*.adoc`). Default off so a symbol mentioned only in CHANGELOG still yields `uncertain`; pass when you want a code-only blast radius (binder channels are unaffected). | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the exact verdict rule (unsafe if strict_refs>0 OR callers>0, uncertain on text-only hits, safe only on all-zero) and the semantics of each channel. It is silent on read-only nature, auth/permissions, and rate limits, keeping it from a full 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then the verdict rule, then the results shape — a logical progression. It is dense and slightly long with heavy parenthetical/backtick notation, but nearly every clause carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description compensates by spelling out the return shape { symbol, verdict, verdict_explanation, channels{...} } and each channel block's fields. Combined with the fully-covered input schema, an agent has everything required to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 12 parameters; baseline is 3. The description references concepts (transitive_callers channel, prose-file text hits) that map to params like `depth` and `exclude_docs`, but adds no syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Delete-safety blast-radius report") and precisely scopes it as composing four named reference channels into a single verdict. An agent can distinguish it from siblings like `usages`, `grep`, and `callers` because the description explicitly names those as the channels it subsumes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use this BEFORE proposing to delete or rename a symbol" gives an explicit trigger, and it routes the agent away from the manual usages → grep → callers dance. This is a clear when-to-use with the alternatives named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
implementationsA
Find every concrete type that extends a base class / implements a trait / interface. Walks the indexed inheritance edges (covers generic-parameterised bases). Prefer over grep for find all subclasses of Foo — grep misses : Foo<T>, indirect inheritance, and trait impls; this resolves the real hierarchy. Supports diff scoping: since / since_branched / changed_only (mutually exclusive) to restrict to recently-touched code.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | DEPRECATED — use `symbol`. Pre-v1.7 alias, still accepted; emits a deprecated_args notice in _meta. | |
| limit | No | Max results | |
| since | No | Restrict results to files changed between `<rev>..HEAD` (accepts anything `git diff` understands: `main`, `HEAD~3`, `origin/main`, SHA). Mutually exclusive with `since_branched` and `changed_only`. | |
| symbol | Yes | Exact name of the base class / trait / interface — canonical key (v1.7+). | |
| exclude | No | Blacklist results by path glob; wins over include (repeatable) | |
| include | No | Whitelist results by path glob, gitignore syntax (repeatable) | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running (default: true) | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| changed_only | No | Restrict results to working-tree changes (staged + unstaged + untracked). Mutually exclusive with `since` and `since_branched`. | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true. | |
| since_branched | No | Restrict results to files changed since this branch diverged from `origin/main` (or `main`/`master`). Mutually exclusive with `since` and `changed_only`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does disclose meaningful behavior: it walks indexed inheritance edges, covers generic-parameterised bases, catches indirect inheritance, and has three mutually exclusive scoping switches. It omits return-shape and error/staleness edge cases, but the core behavioral profile is well conveyed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with zero waste, front-loading the core purpose before the grep comparison and the diff-scoping note. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter tool with no output schema and no annotations, the description leaves gaps: it never indicates what a result looks like (type names, paths, locations) and does not address the `subtypes` sibling. Parameter coverage is complete via the schema, so it is merely adequate rather than thorough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description only reinforces the mutual exclusivity of `since`/`since_branched`/`changed_only`, which the schema already states for each parameter, adding no meaning beyond structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource ('Find every concrete type that extends a base class / implements a trait / interface') and names the scope it walks (indexed inheritance edges). It clearly separates itself from grep, but does not distinguish itself from the sibling tool `subtypes`, which appears to cover similar ground.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-prefer rule ('Prefer over grep for `find all subclasses of Foo`') with concrete reasons grep fails, and explains the three mutually exclusive diff-scoping modes. It stops short of addressing when to choose this over the sibling `subtypes`, leaving one routing question unresolved.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
indexA
Build or rebuild the vex index from scratch. Run once per project; use update afterward for incremental refreshes. Set semantic=true to also generate embeddings (slower; required for find_similar / similar / duplicates).
| Name | Required | Description | Default |
|---|---|---|---|
| gpu | No | Use the GPU for embedding generation if this vex build supports it (DirectML on Windows / CoreML on macOS prebuilts; CUDA via source build), with silent CPU fallback. Only speeds up cold/large semantic builds. Omit to let .vex.toml gpu/device or $VEX_DEVICE decide; pass false to force CPU even when config enables GPU. | |
| device | No | Advanced: pin a specific embedding execution provider (cpu | auto | cuda | directml | coreml). Mutually exclusive with `gpu`. | |
| semantic | No | Also generate per-symbol embeddings (enables semantic search / similar / duplicates; adds ~30-90s on a medium repo) | |
| workspace | No | Multi-repo: fan out across every repo declared in the nearest `.vex-workspace.toml` (set `project_root` at or above it — the manifest is found by walking up). Results become an object `{workspace, repos:[...]}` grouped by repo, NOT the flat per-tool array — branch on shape. `why` is ignored in workspace mode (single-repo only). | |
| project_root | Yes | Absolute path to the project root to index |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose meaningful behavior: the build starts 'from scratch' (implying replacement of any prior index), it is a once-per-project operation, and semantic mode is slower. It stops short of explicitly stating that an existing index is overwritten or that the operation is expensive/irreversible, which would raise this to a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each earning its place: the action, the run-once/use-update rule, and the semantic prerequisite. Front-loaded and free of filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter tool with no output schema and no annotations, the description covers the essential action, lifecycle guidance, and the key semantic prerequisite. Minor gaps remain around workspace-mode fan-out and what happens to pre-existing index data, but structured fields cover the parameter detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (gpu, device, semantic, workspace, project_root) is already documented in detail. The description only reiterates the semantic flag and its downstream effect; it adds no syntax or constraints beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Build or rebuild the vex index from scratch') and explicitly distinguishes itself from the sibling `update` tool by contrasting full build vs incremental refresh. An agent can route between the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance ('Run once per project; use update afterward for incremental refreshes') and names the condition that selects the alternative. It also states the prerequisite for dependent tools (semantic=true required for find_similar / similar / duplicates).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
modulesA
De-facto modules: clusters of symbols that call/reference each other (deterministic Leiden-CPM over call + ref + hierarchy edges, computed on full vex index). Without symbol: list clusters with a label (dominant path prefix; a bare file path when the cluster is a single file), size, cohesion and hub symbols. With symbol: that symbol's cluster and its members (limit caps the matching symbols). Use for what are the modules / which module is X in instead of reading directory listings. Requires a v9 index built by vex index; after vex update clusters are frozen and flagged stale. Returns an empty result with empty_reason and a hint on older indexes or when built with --no-clusters. Cluster ids are stable only within one full-index generation: do not persist them across vex index runs.
| Name | Required | Description | Default |
|---|---|---|---|
| sort | No | Order clusters by in-scope size or by cohesion; ties by cluster id. | size |
| limit | No | Max clusters to list, or max matching symbols when `symbol` is given (per repo with `workspace`). Must be at least 1. | |
| symbol | No | Symbol whose cluster to show. Omit to list all clusters. | |
| exclude | No | Blacklist members by path glob; wins over include (repeatable) | |
| include | No | Whitelist members by path glob, gitignore syntax (repeatable). A cluster is shown iff at least one member is in scope. | |
| members | No | Members to list per cluster, ordered by path then line (default: 0 when listing, 25 for a `symbol` lookup). Must be in `[0, 10000]`. | |
| min_size | No | Hide clusters with fewer in-scope symbols than this. List mode only; ignored when `symbol` is given. Must be in `[1, 1000000]`. | |
| workspace | No | Multi-repo: fan out across every repo declared in the nearest `.vex-workspace.toml` (set `project_root` at or above it — the manifest is found by walking up). Results become an object `{workspace, repos:[...]}` grouped by repo, NOT the flat per-tool array — branch on shape. `why` is ignored in workspace mode (single-repo only). | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running (default: true) | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so: it states the v9 index prerequisite, the stale/frozen behavior after `vex update`, the empty-result contract with `empty_reason` and hints for old or `--no-clusters` indexes, and the critical caveat that cluster ids are stable only within one full-index generation and must not be persisted. This is exactly the behavioral context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core definition and mode split, then prerequisites and caveats. Dense and information-rich with little waste, though a few sentences (id stability, empty_reason) are packed into long clauses that could be split for faster scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter, no-annotation, no-output-schema tool, the description covers purpose, both operating modes, prerequisites, failure/empty behavior, and stability caveats. Nothing essential to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all 13 parameters (baseline 3). The description still adds value by clarifying `limit`'s dual meaning (clusters vs matching symbols), the omit-`symbol`-to-list semantics, and the shape change in `workspace` mode, going beyond the schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('De-facto modules: clusters of symbols that call/reference each other') and even names the algorithm (deterministic Leiden-CPM over call + ref + hierarchy edges). An agent can distinguish this from directory-listing or symbol-lookup siblings without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing: 'Use for `what are the modules` / `which module is X in` instead of reading directory listings.' It also splits behavior by the `symbol` argument (omit to list clusters, provide to get one cluster's members), so the agent knows exactly which mode to pick.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
outlineA
List every symbol (kind + line range) in a single source file via cached tree-sitter parse. Prefer over Read when you only need the file's structure (what's in here?) rather than the full byte stream — outline returns ~50 lines of structured records vs reading thousands of lines of source.
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | DEPRECATED — use `path`. Pre-v1.7 alias, still accepted; emits a deprecated_args notice in _meta. | |
| path | Yes | Filesystem path to the source file — canonical key (v1.7+). Absolute or relative to project_root. | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden and does disclose useful behavior: results come from a cached tree-sitter parse and return roughly 50 structured records versus thousands of source lines. It does not mention language support limits or what happens on unparseable/binary files, leaving a small gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose before the comparison clause. The size/format payoff ('~50 lines vs thousands') justifies its length and nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, but the description compensates by naming the returned fields (kind + line range) and approximate volume, which is enough for an agent to use the result. It stops short of covering supported languages or failure modes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents `path`, `file` (deprecated alias), and `project_root`. The description adds no parameter-level detail beyond the schema, which is the baseline 3 case when structured fields do the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List every symbol ... in a single source file') plus the return shape ('kind + line range') and the mechanism ('cached tree-sitter parse'). It explicitly distinguishes itself from the sibling-like Read tool, so an agent can route without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rule ('Prefer over Read when you only need the file's structure') and contrasts it with the full byte stream alternative. The routing condition is concrete rather than implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pathsA
Enumerate every caller chain from from to to in the persistent call graph (multi-hop, max 6 by default). Prefer over repeated callers calls when you need to know how a function gets reached from a known entry point — paths walks the edges itself in a single response. Requires a v4 index with call graph (built without --no-call-graph).
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | Exact name of the destination function (callee being investigated). | |
| from | Yes | Exact name of the starting function (caller / entry point). | |
| exclude | No | Blacklist intermediate steps by path glob; wins over include (repeatable) | |
| include | No | Whitelist intermediate steps by path glob, gitignore syntax (repeatable) | |
| max_hops | No | Maximum hops between from and to | |
| max_paths | No | Maximum paths to enumerate (caps output, aborts traversal early) | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running (default: true) | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose a meaningful behavioral prerequisite: it requires a v4 index built with call graph, and it notes that paths walks the edges itself in a single response rather than requiring repeated calls. It does not explicitly state read-only safety or describe the return shape, but the prerequisites and execution model are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action, then the sibling comparison, then the prerequisite. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only traversal tool with no output schema, the description covers purpose, the sibling alternative, and the index prerequisite well. It could briefly state what the response contains (the enumerated chains), but the essential call-time information is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every one of the 11 parameters is already documented in the schema (including the auto_update/staleness and exclude_tests nuances). The description only restates the max_hops default of 6, adding no semantics beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (enumerate) and resource (every caller chain from `from` to `to`) with scope qualifiers (multi-hop, max 6 by default). An agent can immediately distinguish this from siblings like callers, callees, and reachable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative ('Prefer over repeated `callers` calls') and the condition that selects it (when you need to know how a function gets reached from a known entry point), plus a prerequisite (v4 index with call graph). It stops short of stating when NOT to use it, e.g. relative to reachable or impact.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
patternA
Structural AST pattern matching: match code by shape, not text. Metavars: $NAME captures an identifier or balanced expression, $_ is a wildcard, $$$ is an anonymous ellipsis, $$$NAME / $$NAME is a named ellipsis that captures multi-line bodies or arg lists, repeated metavars enforce back-reference equality. Composition: space-flanked && and || join sub-patterns (AND requires both shapes in the file with shared captures agreeing; OR takes the union). Prefer over grep / ast-grep for cross-language structural queries — grep cannot match nested syntax, and ast-grep needs per-language scripts; vex pattern works on the cached tree-sitter parse with a skeleton prefilter (~10-50ms). Set why: true to inspect indexed vs live-scan mode. Supports diff scoping: since (rev), since_branched (since this branch diverged from main), changed_only (working-tree changes) — mutually exclusive.
| Name | Required | Description | Default |
|---|---|---|---|
| why | No | Surface a ScanTrace under `_meta.why` in the response: mode (indexed/live_scan), root_kind_inferred, candidate_files / total_files, fallback_reason. | |
| lang | Yes | Language: rust, python, typescript, go, java, csharp, ruby, kotlin, swift, cpp, php, sql, markdown | |
| limit | No | Max matches to return | |
| since | No | Restrict results to files changed between `<rev>..HEAD` (accepts anything `git diff` understands: `main`, `HEAD~3`, `origin/main`, SHA). Mutually exclusive with `since_branched` and `changed_only`. | |
| exclude | No | Blacklist results by path glob; wins over include (repeatable) | |
| include | No | Whitelist results by path glob, gitignore syntax (repeatable) | |
| pattern | Yes | Structural pattern with $METAVARS (e.g. `fn $NAME($$ARGS) -> Result<$T, $E> { $$$BODY }`, `interface $N || class $N`). NOT regex — see grep for regex. | |
| changed_only | No | Restrict results to working-tree changes (staged + unstaged + untracked). Mutually exclusive with `since` and `since_branched`. | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| since_branched | No | Restrict results to files changed since this branch diverged from `origin/main` (or `main`/`master`). Mutually exclusive with `since` and `changed_only`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses metavar semantics, back-reference equality, composition rules, the cached tree-sitter parse with a skeleton prefilter (~10-50ms), and mutual exclusivity of diff-scoping flags. It does not describe the response shape beyond the optional `_meta.why` trace, leaving return-format behavior unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: core concept first, then metavars, composition, sibling comparison, and scoping flags. Every sentence carries information, though the metavar/composition sentences are run-on and could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter tool with no annotations and no output schema, the description covers the operationally critical details: pattern language, performance model, scoping exclusions, and the debug flag. It stops short of describing result structure, but the `why` trace and schema docs fill most of the remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond the schema: metavar grammar ($NAME, $_, $$$, $$$NAME), composition operators with AND/OR capture semantics, and the interaction rules for the diff-scoping parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource: 'Structural AST pattern matching: match code by shape, not text.' It immediately contrasts with grep and ast-grep, so an agent can distinguish it from the sibling grep tool without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternatives and the conditions selecting this tool: 'Prefer over grep / ast-grep for cross-language structural queries — grep cannot match nested syntax, and ast-grep needs per-language scripts.' It also documents mutually exclusive diff-scoping parameters (since / since_branched / changed_only).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reachableA
Every symbol that transitively calls target (the full upstream blast radius). Prefer over repeated callers walks when assessing the impact of changing a function — reachable does the closure in one call. Requires a v4 index with call graph.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results | |
| target | Yes | Exact symbol name whose callers (direct + transitive) you want. | |
| exclude | No | Blacklist results by path glob; wins over include (repeatable) | |
| include | No | Whitelist results by path glob, gitignore syntax (repeatable) | |
| max_hops | No | Maximum hops to walk back from target | |
| workspace | No | Multi-repo: fan out across every repo declared in the nearest `.vex-workspace.toml` (set `project_root` at or above it — the manifest is found by walking up). Results become an object `{workspace, repos:[...]}` grouped by repo, NOT the flat per-tool array — branch on shape. `why` is ignored in workspace mode (single-repo only). | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running (default: true) | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it discloses a real prerequisite: 'Requires a v4 index with call graph', which is not derivable from the schema. It does not, however, cover the workspace-mode shape change, staleness/auto-update semantics, or truncation-by-limit behavior, all of which live only in schema property text.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the definition and scope, then the routing rule, then the prerequisite. No filler and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 11-parameter tool with no output schema and no annotations, the description covers the core semantics and one hard prerequisite. The remaining behavior (workspace result shape, limit truncation, stale-index handling) is fully documented in the schema, so coverage is adequate though not rich.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 11 parameters (including the workspace shape change and exclude_tests path-only caveat). The description adds only the notion of transitive closure, which is really purpose rather than parameter meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Every symbol that transitively calls `target`') and immediately frames the scope as the 'full upstream blast radius', which is unambiguous. It explicitly differentiates itself from the sibling `callers` tool, so an agent can select correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing guidance: 'Prefer over repeated `callers` walks when assessing the impact of changing a function'. This names the alternative, the scenario that selects it, and the reason (closure in one call). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchA
Hybrid structural + semantic code search across the indexed codebase. Fuses FST exact + BM25 + semantic channels in a single ranked list (~4ms FST hit, ~7-15ms with semantic). Prefer over grep for symbol or identifier lookup — grep does a full-scan (seconds on large repos) and returns line matches; this returns ranked symbol records with kind, signature, and line ranges. Use this when you need to find a definition by name, signature shape, or meaning rather than guessing a regex. Supports filter (substring path filter), kind (kind-boost / restrict), context_path (proximity hint), no_bm25 (disable BM25 channel), and no_stale_check (skip pre-call staleness probe).
| Name | Required | Description | Default |
|---|---|---|---|
| why | No | Surface a JSON trace under `_meta.why` in the response: normalized query, per-channel hits (FST/BM25/semantic/fuzzy), filter_applied snapshot | |
| kind | No | Boost results matching one or more kinds (repeatable). Canonical names (function, struct, class, …) plus aliases: def, comment, test, ref. | |
| limit | No | Max results | |
| query | Yes | Free-text query: symbol name, partial name, signature snippet, or natural-language description. Not for regex (use grep) or exact-only resolution (use find_symbol). | |
| since | No | Restrict results to files changed between `<rev>..HEAD` (accepts anything `git diff` understands: `main`, `HEAD~3`, `origin/main`, SHA). Mutually exclusive with `since_branched` and `changed_only`. | |
| exclude | No | Blacklist results by path glob (wins over include); repeat for multiple globs | |
| include | No | Whitelist results by path glob, gitignore syntax (e.g. 'tests/**'); repeat for multiple globs | |
| no_bm25 | No | Disable the BM25 channel for this query (auto-on when the index has BM25 data otherwise). | |
| no_async | No | Exclude async/suspend functions | |
| semantic | No | Enable the semantic vector channel (requires `vex index --semantic`); adds ~3-10ms but lets natural-language queries hit | |
| code_only | No | (v1.20.0 D4) Drop results in prose-format files (`*.md`/`*.markdown`/`*.txt`/`*.rst`/`*.adoc`). Default off so 'README' still finds the README; pass for code-intent queries where CHANGELOG/README headings would pollute the top of the result list. | |
| workspace | No | Multi-repo: fan out across every repo declared in the nearest `.vex-workspace.toml` (set `project_root` at or above it — the manifest is found by walking up). Results become an object `{workspace, repos:[...]}` grouped by repo, NOT the flat per-tool array — branch on shape. `why` is ignored in workspace mode (single-repo only). | |
| async_only | No | Keep only async/suspend functions | |
| visibility | No | Keep only symbols whose signature contains this explicit visibility keyword (no inferred defaults) | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running (default: true) | |
| filter_path | No | Substring path filter applied to result paths (single substring; use include/exclude for glob patterns). Legacy alias: `filter`. | |
| sealed_only | No | Keep only sealed (or Java-`final`) types | |
| static_only | No | Keep only static class members | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| changed_only | No | Restrict results to working-tree changes (staged + unstaged + untracked). Mutually exclusive with `since` and `since_branched`. | |
| context_path | No | Boost results near this file path (e.g. the agent's current editor file). | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true (which already refreshes). | |
| since_branched | No | Restrict results to files changed since this branch diverged from `origin/main` (or `main`/`master`). Mutually exclusive with `since` and `changed_only`. | |
| exclude_generated | No | Drop machine-generated files (protobuf stubs, sqlc output, bindgen bindings, ORM schemas) from the results, recognised from the generator's header banner. Pass on repos that check in generated code, where the stubs outnumber and outrank hand-written symbols. Heuristic: a generator that writes no banner is not detected, so it under-reports rather than hiding hand-written code. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses return shape (ranked symbol records with kind, signature, line ranges) and performance (~4ms FST, ~7-15ms with semantic). It omits notable side-effect behavior such as the default auto_update bootstrapping/refreshing the index and the workspace-mode result-shape change, which live only in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core capability, then routes to alternatives, then summarizes params. Dense and mostly waste-free, though the trailing 'Supports ...' list leans toward enumeration rather than insight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 26-parameter tool with no output schema, the description covers purpose, routing, latency, and return record shape well. Gaps remain around the workspace-mode shape change and auto_update side effects, but the structured schema compensates.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema descriptions are unusually thorough, so the baseline is 3. The description restates filter/kind/context_path/no_bm25/no_stale_check in condensed form, giving a light mental model, but adds little beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Hybrid structural + semantic code search across the indexed codebase') and quantifies the mechanism (FST + BM25 + semantic fused into one ranked list). It also distinguishes itself from grep and find_symbol by name, so an agent can route without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says to prefer this over grep for symbol/identifier lookup and explains why (grep full-scans and returns line matches vs ranked symbol records). It names the alternatives (grep for regex, find_symbol for exact-only resolution) and the condition that selects each.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
showA
Extract the full source body of one or more symbols by name (function, class, struct, etc.) using cached symbol byte-offsets (~4ms per symbol). Prefer over Read when you need a specific definition — show returns just that body, while Read pulls the entire file (often 10-100x more tokens). Accepts an array, so a single call replaces several Read calls. Phase 13.3 truncation: signature_only (signature line only), head (first N body lines), no_body (signature + leading doc only), collapsed (collapse nested methods — v1.9 NO-OP). Also supports filter (substring path filter), kind (kind-restrict), context_path (proximity hint), and no_stale_check.
| Name | Required | Description | Default |
|---|---|---|---|
| head | No | Phase 13.3: print only the first N body lines and append `... (M more lines)`. Mutually exclusive with `signature_only`, `no_body`, `collapsed`. | |
| kind | No | Boost results matching one or more kinds (repeatable). Same vocabulary as `search.kind`. | |
| limit | No | Max bodies returned per symbol name (handles overloads / duplicates) | |
| symbol | No | DEPRECATED — use `symbols: [name]`. Pre-v1.7 singular alias, still accepted; emits a deprecated_args notice in _meta. | |
| exclude | No | Blacklist results by path glob; wins over include (repeatable) | |
| include | No | Whitelist results by path glob, gitignore syntax (repeatable) | |
| no_body | No | Phase 13.3: print signature + leading docstring only; drop the body. Mutually exclusive with `signature_only`, `head`, `collapsed`. | |
| symbols | Yes | Exact symbol names to extract — canonical key (v1.7+). Pass the array form even for a single symbol. | |
| collapsed | No | Phase 13.3: collapse nested methods inside a class/impl/module. v1.9 NO-OP (flag-shape stable; emits a stderr warning). Mutually exclusive with `signature_only`, `head`, `no_body`. | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running (default: true) | |
| filter_path | No | Substring path filter applied to result paths (single substring; use include/exclude for glob patterns). Legacy alias: `filter`. | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| context_path | No | Boost results near this file path (e.g. the agent's current editor file). | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true. | |
| signature_only | No | Phase 13.3: print only the signature line(s). Mutually exclusive with `head`, `no_body`, `collapsed`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does substantial work: it discloses the ~4ms-per-symbol cost model, the 10-100x token advantage, the truncation modes with their mutual-exclusivity rules, and the deprecated-flag behavior surfaced in the schema. It never explicitly states the operation is read-only and non-mutating, which an agent would still want confirmed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the value proposition (extract symbol body, fast, cheaper than Read) before listing the optional knobs. It is dense and mostly waste-free, though the long final sentence stacking several flags reads as a run-on and could be broken up.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-parameter tool with no output schema and no annotations, the description covers usage, cost model, staleness handling, and mode semantics well. It stops short of describing the return shape or what happens when a requested symbol name is not found, which are the remaining gaps an agent would care about.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so a baseline of 3 is warranted, but the description adds real meaning on top: it groups and explains the Phase 13.3 truncation modes (signature_only, head, no_body, collapsed), notes that collapsed is a v1.9 NO-OP, and characterizes filter, kind, context_path, and no_stale_check in prose the schema only states tersely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb and resource — 'Extract the full source body of one or more symbols by name (function, class, struct, etc.)' — with a mechanism (cached symbol byte-offsets) that makes the operation unambiguous. It sharply separates itself from Read, but does not differentiate against the other listed siblings such as find_symbol, outline, or search, so it falls short of the top mark.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use rule: 'Prefer over Read when you need a specific definition.' It also explains the batching rationale (a single array call replaces several Read calls). It offers no when-not-to-use conditions or routing toward find_symbol/outline siblings, so it is clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
similarA
Nearest neighbours of an EXISTING symbol by its stored embedding (HNSW lookup, ~7-15ms). Distinct from find_similar (which embeds a free-text query). Use this when you have a function in hand and want what else in this repo looks like it? — useful for dedup, refactor planning, and finding parallel implementations. Requires vex index --semantic. Supports diff scoping: since (rev), since_branched, changed_only (mutually exclusive) and no_stale_check.
| Name | Required | Description | Default |
|---|---|---|---|
| why | No | Surface a JSON trace under `_meta.why`: seed resolution, applied threshold, candidates before/after path filter, filter snapshot. | |
| name | No | DEPRECATED — use `symbol`. Pre-v1.7 alias, still accepted; emits a deprecated_args notice in _meta. | |
| limit | No | Max results | |
| since | No | Restrict results to files changed between `<rev>..HEAD`. Mutually exclusive with `since_branched` and `changed_only`. | |
| symbol | Yes | Exact name of an existing indexed symbol to use as the seed — canonical key (v1.7+). | |
| exclude | No | Blacklist results by path glob; wins over include (repeatable) | |
| explain | No | Include reasoning per match: identifier-set Jaccard overlap + truncated unified diff between bodies | |
| include | No | Whitelist results by path glob, gitignore syntax (repeatable) | |
| threshold | No | Minimum cosine similarity (0.0..1.0); raise to tighten matches | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running (default: true) | |
| filter_path | No | Substring path filter applied to result paths (single substring; use include/exclude for glob patterns). Legacy alias: `filter`. | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| changed_only | No | Restrict results to working-tree changes (staged + unstaged + untracked). Mutually exclusive with `since` and `since_branched`. | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true. | |
| since_branched | No | Restrict results to files changed since this branch diverged from `origin/main` (or `main`/`master`). Mutually exclusive with `since` and `changed_only`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does reasonably well: it discloses the lookup mechanism (HNSW), expected latency (~7-15ms), the hard prerequisite (semantic index), and the mutual-exclusivity constraint among since/since_branched/changed_only. It does not describe failure modes when the seed symbol is missing or what staleness handling looks like by default beyond the no_stale_check note.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded: purpose, differentiation, use cases, prerequisite, then constraints. Nearly every clause carries information, though the trailing diff-scoping enumeration partly duplicates the schema descriptions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-parameter tool with no annotations and no output schema, the description supplies the operationally critical context (mechanism, latency, prerequisite, flag interactions) an agent needs before calling. Output shape is left unaddressed, but no output schema exists to defer to, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; every parameter is already documented, including the deprecation of `name` and the mutual exclusivity of the diff-scoping flags. The description mostly restates those constraints rather than adding syntax or interaction detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('nearest neighbours of an EXISTING symbol by its stored embedding') and explicitly distinguishes itself from find_similar, which handles free-text queries. The distinction is actionable without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative (find_similar), the condition that selects this tool ('when you have a function in hand'), concrete use cases (dedup, refactor planning, parallel implementations), and the enabling prerequisite ('Requires `vex index --semantic`').
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusA
Report index statistics: symbol count, byte size, embedding presence, last-update timestamp. Use to confirm an index exists and is fresh before running search-shaped tools.
| Name | Required | Description | Default |
|---|---|---|---|
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a passive inspection by listing only statistics, but never states it is side-effect-free, does not say what happens when no index exists (error vs empty stats), and mentions no permissions or cost. Adequate but thin for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with zero waste; the statistics list comes first and the usage guidance second. Front-loaded and appropriately sized for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, enumerating the returned statistics is exactly the right compensation, and it covers the 1-param schema adequately. The only gap is undefined behavior when the index is absent, which an agent would likely want to know for a pre-flight check tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (project_root) and schema coverage is 100%, so the schema already documents its meaning and default. The description adds nothing about path resolution or the working-directory default, making this the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Report index statistics') and enumerates the exact contents (symbol count, byte size, embedding presence, last-update timestamp), so an agent knows precisely what comes back. It does not explicitly contrast itself with the similarly-named 'index'/'update' siblings, but the described payload is distinctive enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear use condition: 'confirm an index exists and is fresh before running search-shaped tools.' This is genuine routing guidance. It stops short of naming alternatives or stating when not to use it (e.g., after running index/update).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
subtypesA
Find every TRANSITIVE subtype of a base class / interface — the full descendant tree via extends/implements edges (not just direct implementations; use implementations for direct-only). Requires a v8+ index with hierarchy edges (no live-walk fallback) — if the index predates this feature or has no hierarchy section, this returns an empty result with a hint to re-run vex index. Supports diff scoping: since / since_branched / changed_only (mutually exclusive) to restrict to recently-touched code.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | DEPRECATED — use `symbol`. Pre-v1.7 alias, still accepted; emits a deprecated_args notice in _meta. | |
| depth | No | Max BFS hops (transitive descent) from the queried type. Bounds how many inheritance levels deep the search goes; independent of the mandatory cycle-detection guard. Must be in `[1, 4096]` (64 is a generous default real hierarchies never approach). | |
| limit | No | Max results | |
| since | No | Restrict results to files changed between `<rev>..HEAD` (accepts anything `git diff` understands: `main`, `HEAD~3`, `origin/main`, SHA). Mutually exclusive with `since_branched` and `changed_only`. | |
| symbol | Yes | Exact name of the base class / trait / interface — canonical key. | |
| exclude | No | Blacklist results by path glob; wins over include (repeatable) | |
| include | No | Whitelist results by path glob, gitignore syntax (repeatable) | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running (default: true) | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| changed_only | No | Restrict results to working-tree changes (staged + unstaged + untracked). Mutually exclusive with `since` and `since_branched`. | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true. | |
| since_branched | No | Restrict results to files changed since this branch diverged from `origin/main` (or `main`/`master`). Mutually exclusive with `since` and `changed_only`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the v8+ index requirement, the absence of a live-walk fallback, and the failure mode (empty result plus a re-index hint) if the hierarchy section is missing. It does not describe the return shape, but the failure-mode and precondition disclosure is genuinely valuable context beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core purpose, then preconditions, then diff scoping. Every clause earns its place with no redundancy or padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-parameter tool with no annotations and no output schema, the description covers purpose, the critical index precondition, the failure mode, and diff scoping. It leaves the return format undocumented (no output schema exists), which is the main remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so every parameter is already fully documented in the schema; the baseline is 3. The description reinforces the mutual exclusivity of `since`/`since_branched`/`changed_only` and the transitive nature of `depth`, but adds little that the schema descriptions do not already state.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Find every TRANSITIVE subtype of a base class / interface') and immediately scopes it ('full descendant tree via extends/implements edges'), explicitly distinguishing it from the sibling `implementations` for direct-only queries. An agent can route between the two without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative explicitly ('use `implementations` for direct-only') and states the precondition for use (requires a v8+ index with hierarchy edges). It also notes when the diff-scoping params apply. Nothing about tool selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tests_forA
Find test functions that transitively cover a target symbol (Phase 13.10). Walks the call graph backwards from <target>, keeps rows under recognized test-path globs (Rust / Python / TS-JS / Go / Java / Kotlin / C# / C++), stamps each row with a framework label (pytest, jest, go-test, …) so an agent can pick the right runner without parsing paths. Prefer over grep test.*Foo — that misses transitively-covered helpers and produces lots of false positives. v1.20.0 (D5) surface — the CLI subcommand exists since v1.19.0 but was MCP-invisible.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max results to return. | |
| symbol | No | DEPRECATED alias for `target`; still accepted, emits a deprecated_args notice in _meta. | |
| target | Yes | Symbol whose test coverage to find — the function/method/class you want to know is tested. | |
| exclude | No | Blacklist results by path glob; wins over include (repeatable) | |
| include | No | Whitelist results by path glob, gitignore syntax (repeatable) | |
| max_hops | No | Maximum reverse-call-graph hops from `target`. Default 6. | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running (default: true) | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| test_pattern | No | Glob patterns for test paths (repeatable). When set, REPLACES the default pattern set (does NOT append) — pass the full set you want. | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. | |
| include_fixtures | No | Admit non-test-named helpers (fixtures) under test paths via a one-hop forward callee walk. Default off — only `test_*` / `*Test` names surface. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the behavioral load — and it does substantial work: it explains that rows are stamped with a `framework` label for runner selection, that the walk is transitive with a bounded hop count, and that the CLI surface predates the MCP surface. It doesn't cover return format or what stale-index behavior looks like at the result level, but the framework-label disclosure is valuable behavioral context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then mechanism, then the grep contrast, then a version-history note. The version sentence is somewhat meta and could be trimmed, but it's a single trailing sentence and doesn't obscure the main payload.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 12 parameters with full schema coverage but no output schema, the description correctly focuses on what the tool returns conceptually (test rows with framework labels) and the transitive mechanism. It leaves return-format details to the agent's expectations, which is acceptable when no output schema exists, though a note on result shape would have strengthened it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all 12 parameters thoroughly, including the deprecated `symbol` alias, `test_pattern` replace-not-append semantics, and the `async_update`/`_meta.vex.dev/stale` contract. The description adds meaning by explaining WHY the framework label exists (so agents can pick a runner without path parsing), which goes beyond structured data, but it doesn't re-explain the parameters themselves — the schema does that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (find) and resource (test functions that transitively cover a target symbol) with the exact mechanism (backward call-graph walk). Distinguishes itself from the sibling `callers`/`callees` by being test-path-filtered and from `grep` by reason. An agent can immediately tell what it returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative (`grep test.*Foo`) and the reason it's wrong (misses transitive helpers, false positives), giving a clear when-to-use-this-not-that rule. This is the strongest kind of routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
updateA
Incremental index refresh: only re-parses files whose mtime changed since the last index. Prefer over index when an index already exists — typically <1s on small change sets vs full rebuild cost. Most other tools default to auto_update=true and call this implicitly.
| Name | Required | Description | Default |
|---|---|---|---|
| gpu | No | Use the GPU for embedding generation if this vex build supports it, with silent CPU fallback. Mostly a no-op for incremental updates (few/zero embeddings recomputed). Omit to let .vex.toml gpu/device or $VEX_DEVICE decide; pass false to force CPU. | |
| device | No | Advanced: pin a specific embedding execution provider (cpu | auto | cuda | directml | coreml). Mutually exclusive with `gpu`. | |
| semantic | No | Also refresh embeddings for changed files | |
| workspace | No | Multi-repo: fan out across every repo declared in the nearest `.vex-workspace.toml` (set `project_root` at or above it — the manifest is found by walking up). Results become an object `{workspace, repos:[...]}` grouped by repo, NOT the flat per-tool array — branch on shape. `why` is ignored in workspace mode (single-repo only). | |
| project_root | Yes | Absolute path to the project root whose index should be refreshed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the mtime-based incremental mechanism, relative cost, and implicit invocation by other tools. It does not say what happens when no index exists (error vs fallback to full build) or whether the index file is mutated in place, which are the remaining behavioral gaps for a mutating refresh tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: purpose first, then the routing decision, then the implicit-call caveat. Every sentence earns its place with zero padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating refresh tool with no output schema and fully documented parameters, the description covers mechanism, cost, and sibling routing. The main omission is the failure/fallback path when no prior index exists, which an agent would want before invoking it on an unindexed project.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters (gpu, device, semantic, workspace, project_root) in detail. The description adds no parameter-level meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Incremental index refresh') plus the exact mechanism (re-parses only files whose mtime changed). It explicitly contrasts itself with the sibling `index`, so an agent can distinguish the two without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use rule: 'Prefer over `index` when an index already exists', backed by a concrete cost comparison (<1s vs full rebuild). It also warns that most sibling tools call this implicitly via auto_update=true, which is exactly the routing context an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
usagesA
Find every reference to a symbol across the codebase. Prefer over grep for refactor-style find all callers queries — grep on a common identifier returns string-literal and comment noise; usages with strict=true uses the scope-binder to resolve real cross-file refs (Rust/TypeScript/Python/C#/C++). Without strict, runs the legacy refs FST (~4ms) — v1.20.0 also strips the row at the symbol's own definition line and prose mentions in *.md/*.markdown/*.txt/*.rst/*.adoc (override with include_self / include_docs).
| Name | Required | Description | Default |
|---|---|---|---|
| why | No | Surface a JSON trace under `_meta.why`: mode (strict/fst_lookup), mode_legacy (back-compat alias for v1.9.x consumers, removed in v1.12), hits before/after path filter, prefix-suggestion count when no exact hits, def_site_dropped / docs_dropped counts (v1.20.0), filter snapshot. | |
| name | No | DEPRECATED — use `symbol`. Pre-v1.7 alias, still accepted; emits a deprecated_args notice in _meta. | |
| limit | No | Max results | |
| since | No | Restrict results to files changed between `<rev>..HEAD` (accepts anything `git diff` understands: `main`, `HEAD~3`, `origin/main`, SHA). Mutually exclusive with `since_branched` and `changed_only`. | |
| strict | No | Use scope-resolved (type-aware) references from the binder — drops string-literal/comment/wrong-scope noise. Recommended for refactor work; falls back to legacy refs FST on languages without binder support. | |
| symbol | Yes | Exact symbol name to find references to — canonical key (v1.7+). | |
| exclude | No | Blacklist results by path glob; wins over include (repeatable) | |
| include | No | Whitelist results by path glob, gitignore syntax (repeatable) | |
| workspace | No | Multi-repo: fan out across every repo declared in the nearest `.vex-workspace.toml` (set `project_root` at or above it — the manifest is found by walking up). Results become an object `{workspace, repos:[...]}` grouped by repo, NOT the flat per-tool array — branch on shape. `why` is ignored in workspace mode (single-repo only). | |
| auto_update | No | Auto-update the index if stale, or bootstrap it if missing, before running (default: true) | |
| filter_path | No | Substring path filter applied to result paths (single substring; use include/exclude for glob patterns). Legacy alias: `filter`. | |
| async_update | No | With auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false) | |
| changed_only | No | Restrict results to working-tree changes (staged + unstaged + untracked). Mutually exclusive with `since` and `since_branched`. | |
| include_docs | No | (non-strict only) Keep matches in `*.md` / `*.markdown` / `*.txt` / `*.rst` / `*.adoc` files. v1.20.0+ strips them by default — README/CHANGELOG mentions of a symbol are prose, not callers. No-op when `strict=true`. | |
| include_self | No | (non-strict only) Keep the row at the symbol's own definition line. v1.20.0+ strips it by default — `find all callers` queries don't want the declaration showing up as a usage. No-op when `strict=true` (the scope-binder excludes the def-site by construction). | |
| project_root | No | Absolute path to the project root (defaults to the MCP working directory) | |
| exclude_tests | No | Drop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded. | |
| no_stale_check | No | Skip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true. | |
| since_branched | No | Restrict results to files changed since this branch diverged from `origin/main` (or `main`/`master`). Mutually exclusive with `since` and `changed_only`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and discloses mode differences, language support, legacy FST speed (~4ms), fallback behavior, and v1.20.0 default stripping of definition-line/prose matches with overrides. It still omits normal return shape, pagination, and error behavior, so it is strong but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first sentence, and the rest gives relevant mode and filtering context. The single dense paragraph is mostly earned, though version-specific detail like v1.20.0 adds some weight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 19-parameter tool with no output schema and no annotations, the description covers the main behavior, alternatives, and mode selection well. It leaves the normal result shape and pagination undescribed, but the parameter schema is thorough and carries most remaining detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds useful parameter-level meaning beyond the schema by listing which languages strict=true supports, noting the ~4ms non-strict path, and consolidating the include_self/include_docs override behavior, though it does not cover every parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Find every reference to a symbol across the codebase.' It distinguishes itself from sibling grep by naming the refactor-style 'find all callers' use case and explaining grep's string-literal/comment noise.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Prefer over grep for refactor-style `find all callers` queries,' giving the alternative tool and the condition for choosing this one. It also explains strict=true versus non-strict legacy FST behavior, so the agent knows which mode fits the query.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
28 tool updates
v1.27.3- First observed
bundle - First observed
callees - First observed
callers - First observed
capabilities - First observed
check - First observed
diff - First observed
duplicates - First observed
eval - First observed
find_similar - First observed
find_symbol - First observed
grep - First observed
history - First observed
impact - First observed
implementations - First observed
index - First observed
modules - First observed
outline - First observed
paths - First observed
pattern - First observed
reachable - First observed
search - First observed
show - First observed
similar - First observed
status - First observed
subtypes - First observed
tests_for - First observed
update - First observed
usages
TDQS
Scored across 28 tools
Most tools have clearly distinct query modalities, and descriptions add explicit 'prefer over X' guidance for near neighbours. However, the search family (search/find_symbol/find_similar/similar/grep) and the reference family (callers/usages/paths/reachable/impact/bundle) overlap enough that an agent must read the long descriptions to choose correctly.
Names are consistently lowercase snake_case with no camelCase or casing drift. Minor deviation: a few tools use verb prefixes (find_symbol, find_similar, tests_for) while most are bare nouns or bare verbs.
28 tools is heavy for a single MCP server, even a broad code-intelligence one. Many tools are genuinely distinct, but the surface could likely be consolidated (bundle already subsumes show+callers+callees+similar, and search/find_symbol/find_similar/similar form a large cluster).
The surface covers index lifecycle (index/update/status/capabilities/eval), symbol and structural search, call graph traversal, hierarchy, references, impact analysis, tests, history, diffs, modules, duplicates, and semantic similarity. For a read-only code-intelligence agent, there are no obvious dead ends.
Maintenance
Related MCP Connectors
Code intelligence for coding agents: semantic, AST, graph, and full-text search. 279+ languages.
Code intelligence platform for AI agents. 20 tools for architecture, security & impact analysis.
AI-powered codebase analysis — call graphs, security, dead code, complexity. 150+ tools.
Search GitHub, npm, PyPI, StackOverflow, ArXiv from one MCP — built for coding agents.
Related MCP Servers
- AlicenseAqualityBmaintenanceEnables AI assistants to search and analyze codebases using Abstract Syntax Tree (AST) pattern matching with ast-grep. Supports structural code search, pattern testing, and AST visualization across multiple programming languages.4471MIT
- AlicenseAqualityAmaintenanceKnowledge graph for token-efficient code reviews. Builds a structural map of your codebase with Tree-sitter, tracks changes incrementally, and gives AI agents precise context via MCP tools. Features fixed multi-word search, qualified call resolution, dual-mode embedding (ONNX local + LiteLLM cloud), and output pagination.668Apache 2.0
- AlicenseNot gradedqualityNot gradedmaintenanceCodeGraph — Open-source code intelligence MCP server. Builds a semantic graph of your codebase (functions, classes, imports, call chains) and exposes it through 31 tools. Callers, callees, impact analysis, complexity metrics, unused code detection, AI context assembly, persistent memory, cross-project search. 15 languages via tree-sitter. Single Rust binary, local-first.301 npm-
- FlicenseBqualityBmaintenanceLocal-first code intelligence for AI coding assistants: MCP tools, symbol graph search, impact analysis, and auto-index watch for Cursor, Claude Code, and Codex.622-