Skip to main content
Glama

duplicates

Find similar code symbols repo-wide by comparing embeddings to plan refactoring and deduplicate. Supports diff scoping and filtering.

Instructions

Repo-wide near-duplicate scan: pairs of symbols whose embeddings exceed threshold. Use for refactor planning (where else does this logic live?) and dedup. Prefer over manual similar-walks — duplicates evaluates all pairs once with min_body_lines filtering out trivial bodies. Requires vex index --semantic. Supports diff scoping: since (rev), since_branched, changed_only (mutually exclusive) and no_stale_check.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
whyNoSurface a JSON trace under `_meta.why`: applied threshold + min_body_lines, pairs before/after path filter, filter snapshot.
limitNoMax pairs to return
sinceNoRestrict pairs to files changed between `<rev>..HEAD`. Mutually exclusive with `since_branched` and `changed_only`.
excludeNoBlacklist pairs by path glob — a pair is dropped when either side matches (repeatable)
explainNoInclude reasoning per pair: identifier-set Jaccard overlap + truncated unified diff between the two bodies
includeNoWhitelist pairs by path glob — a pair is kept when at least one side matches (repeatable)
thresholdNoMinimum cosine similarity (0.0..1.0); 0.9 keeps only very close pairs
auto_updateNoAuto-update the index if stale, or bootstrap it if missing, before running (default: true)
filter_pathNoSubstring path filter — keep pairs where at least one symbol's path contains this substring. Legacy alias: `filter`.
async_updateNoWith auto_update, refresh a stale index in the background instead of waiting for it: results come from the index already on disk and _meta.vex.dev/stale says so (default: false)
changed_onlyNoRestrict pairs to working-tree changes (staged + unstaged + untracked). Mutually exclusive with `since` and `since_branched`.
project_rootNoAbsolute path to the project root (defaults to the MCP working directory)
exclude_testsNoDrop test files from the results (tests/ dirs, *_test.*, test_*.py, *.spec.ts, __tests__/, tests.rs, ...; same set as tests_for). Composes with include/exclude. Path-based only: Rust unit tests inside a `#[cfg(test)] mod tests` block of a non-test file are not excluded.
min_body_linesNoSkip symbols with body shorter than this many lines (filters trivial wrappers)
no_stale_checkNoSkip the staleness check that runs before each call; assumes the index is fresh. Redundant when `auto_update` is true.
since_branchedNoRestrict pairs to files changed since this branch diverged from `origin/main` (or `main`/`master`). Mutually exclusive with `since` and `changed_only`.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.27.3

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden well: it discloses the hard prerequisite ('Requires `vex index --semantic`'), the filtering behavior of min_body_lines, and the mutual exclusivity of the three diff-scoping flags. It stops short of describing the return shape (pair list) or how many pairs come back by default.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense, front-loaded sentences: purpose, usage, then prerequisites and scoping flags. Every clause carries information, though the final scoping sentence packs several flags together and reads a bit like a spec dump.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 16-parameter, zero-required, annotation-free tool with no output schema, the description covers the prerequisite, the filtering model, and the diff-scoping exclusivity — the things an agent most needs. The missing piece is what a result actually looks like (pair structure/ordering), which nothing else supplies.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds cross-parameter semantics the schema cannot express — the mutual exclusivity of `since`/`since_branched`/`changed_only` and the fact that `no_stale_check` is redundant under `auto_update`. It restates threshold/min_body_lines meaning rather than extending it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource+scope: 'Repo-wide near-duplicate scan: pairs of symbols whose embeddings exceed threshold.' That clearly separates it from symbol-lookup siblings, though it never names the closest siblings (find_similar, similar) and only alludes to them as 'manual similar-walks.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete use cases ('refactor planning', 'dedup') and a comparison condition ('Prefer over manual similar-walks — duplicates evaluates all pairs once with min_body_lines filtering'). No explicit when-not guidance and no direct routing to find_similar/similar for the single-symbol case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.