pasr
PASR is a local MCP server for finding source and selecting budgeted, source-traceable context in a workspace.
find_files— rank workspace paths by query-term matches in names; can list candidates with an empty query and optional include scopes.find_symbols— find exact/partial symbol definitions asfile:lineranges in Python, JS/TS, and Rust, with include globs and kind filters.find_evidence— search file contents for query terms, returning rankedfile:linehits and bounded read ranges.find_usages— find lexical references to a symbol, returning source lines, enclosing definitions, and read-range hints.select_context— select source spans for a query from files,path:start-endranges, and/or include globs, capped at 1,500 tokens by default; supports outline/definition index, advanced options, and continuation/fingerprint metadata for range reads.All location tools use workspace-relative
includeglobs (*.pyroot-only,**/*.pyrecursive) and return paths/positions or read hints.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@pasrfind where select_context enforces the token budget, include file:line provenance"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
The problem
Agents need enough source to answer correctly, while every retrieved observation can be carried into later model requests. Native grep and bounded file reads already provide useful source and provenance. PASR must justify its additional retrieval, selection, and tool-catalog complexity against that baseline—not against dumping the entire repository.
Related MCP server: token-context-mcp
What PASR does
Budgeted
Live select_context budgets the returned context, including labels, headings,
separators, and redaction—not the MCP JSON envelope or cumulative conversation.
Whole computed spans are admitted only when that final representation fits; spans
are not necessarily complete functions. Lossless selection uses the exact rendered
cost, and optional headers cannot displace source that already fits.
Traceable
Selection receipts record each span's file:line, token count, retrieval signals,
and why it was kept or dropped. Receipts are saved under .pasr/receipts; persistence
is best-effort, and the response reports whether a receipt was written. They cover
PASR-delivered context, not everything an agent reads or does.
Honest
Receipts expose heuristic query classifications, keyword coverage, and routing advice. These are not calibrated correctness probabilities or proof that all necessary evidence was retrieved. Advice must not replace checking the source.
Local by default
No daemon, vector database, or manual index setup is required. Offline selection needs installed dependencies and cached tokenizer data; optional model weights are separate. The client can still forward returned source to a cloud model. Default redaction is a no-op, not automatic secret detection.
Install
Install version 0.4.0 from PyPI.
Install uv and make uvx available on your PATH;
use Python 3.10+ or allow uv to provision it. No PASR checkout is required.
Replace /absolute/path/to/project with the repository to inspect, then choose
the command for your client:
claude mcp add pasr -- uvx --from pasr-mcp==0.4.0 pasr-mcp --workspace /absolute/path/to/project
codex mcp add pasr -- uvx --from pasr-mcp==0.4.0 pasr-mcp --workspace /absolute/path/to/projectAlternatively, merge this MCP configuration into your client's config. Update any
existing pasr entry rather than registering the same server twice:
{ "mcpServers": { "pasr": { "command": "uvx", "args": ["--from", "pasr-mcp==0.4.0", "pasr-mcp", "--workspace", "/absolute/path/to/project"] } } }For the separate v0.4.0 GitHub release wheel instead, use:
python -m pip install https://github.com/Apheironn/pasr/releases/download/v0.4.0/pasr_mcp-0.4.0-py3-none-any.whlAfter installing that wheel, configure your client to run the installed pasr-mcp
executable with --workspace /absolute/path/to/project instead of the uvx launch.
Developers only — source checkout: replace pasr-mcp==0.4.0 after --from
in your existing registration with /absolute/path/to/pasr. For example:
uvx --from /absolute/path/to/pasr pasr-mcp --workspace /absolute/path/to/projectUse this instead of the published-package launch, not a second pasr registration.
It runs that checkout, which may differ from the release; editing it does not update
an existing global installation.
After installation, PASR's tools are available alongside the agent's native tools. The model decides when to search, select context, or read files directly. Smaller selected context does not by itself establish cheaper or more accurate answers.
Per-client setup notes: Claude Code · Cursor · Windsurf.
Without an agent: run uvx --from pasr-mcp==0.4.0 pasr explain "<question>"
from the project to inspect to print a selection receipt.
Five historical CLI examples against pinned public repos are in
examples/.
0.4.0 diagnostics
Version 0.4.0 includes pasr doctor. To diagnose this package installation,
use the same version as your registered server:
uvx --from pasr-mcp==0.4.0 pasr --workspace /absolute/path/to/project doctor --json --timeout 30Omit --json for a readable report. Exit codes are 0 for success, 1 for a
failed diagnostic, and 2 for invalid options. The protocol deadline must be a
finite positive number; subprocess shutdown can take a few additional seconds.
The command checks that the workspace is a directory, then starts a real MCP
server in a disposable fixture to inspect its catalog and select known source.
It never scans or writes the project or its .pasr/ state. Reports contain runtime
versions and check outcomes, not source, absolute paths, environment values or raw
exception messages. It makes no paid model call or automatic report upload;
initial setup/tokenizer encoding downloads can require network access.
A pass does not verify your client's registration, approval settings or answers.
Optional public reports: installation trouble or real-task feedback. Review before posting; do not include secrets, private source or raw transcripts. Doctor is optional; you can report without running it.
Verified Codex onboarding, with explicit limits
On 2026-10-04, Codex CLI 0.160.0 + GPT-6 Luna used an installed 0.3.0 wheel to answer a source-reading question in an isolated workspace containing one unmodified PASR source file. The final recipe used three model-selected PASR calls, 364 selected-context tokens, and source-checked relative line citations. This is a controlled integration smoke, not independent user adoption or evidence of lower total model cost. Setup failures and earlier model turns are retained in the integration record.
For interactive use, approve the known PASR server when the client prompts.
In unattended Codex runs, approval_policy = "never" does not grant MCP tool
permission: the initial model call was blocked. Only for a reviewed, trusted
server/workspace, set default_tools_approval_mode = "approve" inside the existing
[mcp_servers.pasr] table. This is explicit preapproval of that server, not a reason
to disable global safeguards. PASR may write receipts/cache in its workspace;
the host's read-only shell sandbox does not sandbox the MCP process.
Ask for workspace-relative file:line ranges copied from PASR output, as plain
text rather than invented absolute links. The earlier answer added a nonexistent
/workspace prefix; the revised citation instruction produced matching references.
When scripting codex exec, close unused stdin; otherwise it can wait for
additional prompt input. None of these observations proves another client/model
will behave identically.
To repeat this controlled task, build a wheel and run the checked-in driver from
the checkout root. Install Codex first and supply OPENAI_API_KEY through your
normal secret mechanism; this makes paid API requests.
python -m build --outdir dist/demo
python scripts/demo_client.py --codex /absolute/path/to/codex --wheel dist/demo/pasr_mcp-0.4.0-py3-none-any.whl --output eval/agent_bench/results/client-demoOn Windows, pass the actual codex.exe, not its shell wrapper. The output directory
must be new. The driver installs the wheel into a temporary environment, copies
only src/pasr/source_text.py, isolates Codex configuration, and retains the answer,
tool calls, usage, warnings and failures before cleanup. It does not grade answer
correctness, and only the exact supplied API key is redacted: inspect evidence
before sharing. The maintained driver was also exercised with a real model;
see its repeat record.
In one call
One localized question — "how are redirects resolved and followed" — against
psf/requests (verbatim transcript):
Historical CLI selection | Observation |
source-context tokens | 2,718 selected from 42,768 supplied tokens |
provenance |
|
heuristic diagnostic |
|
observed selection time | ~0.3 s offline on this repository |
This is a source-selection illustration, not a native grep/read baseline, an answer-quality result, or the current MCP default cap of 1,500 context tokens.
Evidence and alternatives
Native grep and bounded reads are useful defaults. PASR adds explicit context budgets and selection receipts; that does not establish better answers or a lower total model bill. Aider RepoMap provides structural navigation, and Repomix packages source for a model. These are overlapping workflows, not interchangeable products.
Fresh workflow study (2026-10-05): no pipeline promotion. All 10 PASR methods were measured in 420 local calls. The answering study retained 664 trajectories, including a 400-attempt, 40-question comparison with actual Repomix and Serena components; actual Aider RepoMap also ran in development.
The selected prompt policy used 18.26% more cumulative provider tokens than unchanged split-4 on Luna (paired 95% ratio interval 1.04594–1.35231). Full rubric/source support rose from 3/40 to 4/40 and material-error answers fell from 5 to 3, but the cost objective failed. These stringent completeness scores are not ordinary answer accuracy. Two provider failures retain unknown usage; the user-approved continuation finished the remaining schedule without retries, not as an unamended confirmation. Production retrieval defaults remain unchanged.
See the method-by-method and competitor report, machine-readable summary and licensed evidence bundle. The answering API's conservative cost bound is $0.644405; authoring/coding/review assistant-session tokens are excluded and unmeasured.
Earlier source-reviewed studies (2026-10-03): the combined quality/cost gates
also failed. Their experimental split reader used search_code / read_code
with a host-enforced four-call limit and conservative stopping instructions.
That workflow is not the default installed product:
Study | Supported answers (native / split) | Material errors (native / split) | Token finding and decision |
Previously untouched 40-question confirmation | 28/40 / 28/40 | 3 / 7 | Split used 25.3% fewer cumulative provider tokens; primary gate failed |
Exploratory reuse of the same 40 questions, GPT-6 Luna | 26/40 / 32/40 | 2 / 2 | Mean tokens: native 12,430.875 known subtotal, split 10,698.075; combined gate failed |
The exploratory native run has an interrupted request with unknown usage; its known subtotal is not a complete token total. The same study used actual upstream Aider RepoMap and Repomix context generation under a shared answering harness, not the full Aider coding agent. Results vary by model, and uncertainty does not support equivalent accuracy or universal superiority over native tools or competitors.
See the confirmation record, exploratory matrix, competitor methods and caveats, and pipeline audit, sections 23–24. Earlier selector evaluations and offline localization proxies use different protocols; they are not current end-to-end product wins.
How it works
query + files / line ranges / globs
→ discover safe files, read requested sections, tokenize, line-aligned chunks
→ lossless-under-budget check: does the whole thing already fit? return it
→ range-only body read: ordered whole-line prefix + explicit remaining ranges
→ otherwise, ranked selection:
→ candidates: BM25 + lexical coverage + Python AST / JS/TS/Rust tree-sitter symbols
+ optional hashing / MiniLM scorer
→ fuse unique span ranks per signal score(s) = Σ_r 1 / (k + rank_r(s)), k = 60
→ optional single-source prefix/tail reserve (off in default MCP selection)
→ greedy source-span packing with exact final-context cost ≤ budget
→ classify the query, score confidence, write the receipt
→ return spans + provenance + token accounting + adviceRRF combines rank orders rather than calibrating raw scores. Packing is greedy
(coverage-aware or score-only), not an exact knapsack optimizer. Selected spans
are dependency-ordered before their rendered cost is checked.
Coverage-aware selection preserves literally requested, case-sensitive definitions
before the fused-candidate cutoff and prices complete bodies against the actual
rendered budget. Compatible overlapping source spans are unioned once instead of
discarding the uncovered parts. Remaining choices favor marginal keyword coverage
per token, then source diversity and rank.
BM25 and lexical selection retain exact identifiers and also match dotted,
hyphenated and snake-case components, so needle can match needle_worker inside
raw chunks without treating it as a literal definition-name request. Component
matching does not imply synonym understanding or complete mechanism coverage.
MCP tools
Locators return workspace paths and, where available, line positions or read_lines
hints. select_context accepts paths and path:start-end ranges. A location hint
does not prove that a narrow range contains every mechanism needed for the answer.
Embedders can publish just the two locators with
create_server(root, expose=("find_symbols", "find_evidence")) alongside a host's
native source reader. Their navigation advice does not call unexposed PASR tools.
This is an experimental configuration, not a demonstrated accuracy/token win;
the default five-tool catalog is unchanged.
The equivalent stdio configuration is
uvx --from pasr-mcp==0.4.0 pasr-mcp --workspace /absolute/path/to/project --tools find_symbols,find_evidence.
Keep the host's native grep and source reader available; this catalog does not
contain a source-reading tool.
include paths and globs are relative to the workspace root. *.py searches only
root-level files; **/*.py includes nested Python files. Omit include rather than
guessing the layout. Locator descriptions state this distinction; a matching but
too-narrow scope is not automatically widened.
The first five tools are exposed by default; the others are opt-in.
Tool | Purpose |
| repository-wide content search: explicitly qualified definitions first, then blended rarity, sub-word similarity and reference rank; top hits carry bounded |
| rank files by path/filename match |
| where a symbol is defined, as |
| where a symbol is used: matching lines, enclosing definitions, and bounded |
| budgeted, provenance-tracked slice — |
| discover and select from the top three matching files using only |
| select from required, non-empty known |
| bounded static name-reference approximation; |
| return a prior receipt if it was persisted and is still available |
| increase a prior selection's budget without widening its source ranges or changing outline mode |
Literal symbol names and single filename/path queries retain stopword components:
Where and where.py remain searchable rather than disappearing as prose words.
An exact literal symbol name suppresses partial namesakes; its underscore components
remain available for fallback only when no exact definition matches.
Only a blank or whitespace-only find_files query requests an unranked listing.
Symbol-kind aliases are applied to both the requested filter and discovered kinds,
so a native Python class remains findable with kinds=["class"].
For discovery, an explicit qualified name such as hooks.enforce or
Controller.dispatch prioritizes the matching definition's file over mere mentions.
The qualifier must match the module path or actual enclosing definitions;
this does not resolve import aliases. Broad-query scoring, explicit scopes and
context budgets are unchanged. The opt-in search_code uses the same discovery
ordering before selecting from three files.
Join the path in a hit's provenance with its read_lines and pass that to
select_context(query=..., files=["path.rs:42-56"]) rather than reading the whole file. A suggestion contains the enclosing function when it
is at most 40 lines; otherwise it contains up to eight lines on each side of the
hit. Large functions may require a wider explicit range or find_symbols to locate
the complete definition. Existing snippets, ranking, counts and warnings are retained.
For an unknown location, call search_code(query="cert_verify") to discover real
workspace-relative paths. Discovery accepts no files, include, or other scope
argument; do not guess paths from package names.
For a known location, call the separate opt-in reader:
read_code(query="validation exit state", files=["src/attr/validators.py:73-88"]).
Its files argument is required and non-empty, using the same path/range syntax
as select_context. Empty, missing, directory, mixed-missing, and escaping scopes
fail rather than widening the search. Both tools reject unknown arguments.
Migration: replace search_code(query=..., files=...) with read_code(...);
the old scoped discovery call is rejected, not silently treated as discovery.
Range-only replies include continuation and same-read source_fingerprint
values. Pass remaining ranges to read_code until continuation.files is empty;
do not submit the empty list. Stop if blocked is true. A non-empty query is
still required; explicit ranges determine which lines are read. Plain-file or
mixed scopes remain query-ranked, not sequential whole-file reads.
Scope errors do not trigger automatic discovery, path correction, or retries.
Each call is independent; compare shared-file fingerprints before joining pages.
Publish the compact pair with
uvx --from pasr-mcp==0.4.0 pasr-mcp --workspace /absolute/path/to/project --tools search_code,read_code.
The default five-tool catalog and select_context behavior are unchanged.
This flag does not impose the experiments' four-call limit or stopping policy.
Unlike select_context, the compact pair does not persist selection receipts or
append usage-ledger entries. It returns raw source text with an optional one-line
JSON metadata header, plus the full object in MCP structuredContent.
Discovery charges JSON string quoting/escapes against its context cap; returned
source stays unescaped text in the structured channel. The reader retains the
previous explicit-follow-up pricing and continuation semantics. Advice, metadata,
tool catalogs, and repeated conversation history still cost extra. This interface
separation is not a demonstrated answer-accuracy or cumulative-provider-token win.
Multiple ranges from the same file are unioned: overlapping lines are returned only
once, and gaps stay excluded. An explicitly listed whole file overrides its ranges.
Adding an include pattern does not widen an explicitly ranged file. Ranges also
constrain symbol candidates, outlines, maps and embedded dependency traces.
Complete range reads mean the requested sections are included, not that caller or
dependency behavior has been covered. Follow those relationships when the question
requires them. expand_context increases the budget inside the same scope; request
wider ranges explicitly to read surrounding code.
Range-only body reads are sequential, not query-ranked. When every resolved file
has a range and outline=false, PASR returns a prefix in file-request order, with
merged ranges in ascending line order. It never skips an over-budget line to select
a later match. Whole-file or mixed whole-file/range requests retain ranked selection.
Ranking/window settings do not change range-only body order; optional map/trace
headers still consume budget and can repeat source separately from that body.
The response includes continuation, for example:
{"files": ["path.rs:57-120"], "blocked": false}To advance, pass continuation.files as the next call's files, without include.
An empty list means the extant requested ranges are exhausted, not that the answer
is complete. Ranges are clipped to the current file's end. blocked=true means no
body line advanced: increase the budget where possible or use a direct reader.
The MCP selection cap remains 1,500 tokens; repeating a blocked request cannot help.
Responses carry source_fingerprint; compare shared-file identities before combining
pages. Source edits require a fresh read, not trusting old line coordinates.
Receipts and saved packs retain continuation metadata, but do not freeze future reads.
Every selection is independent. Repeating a request returns the requested source again; PASR does not assume that a previous response remains in the model's context. There is no server-lifetime call ceiling, novelty refusal, or automatic continuation through previously unread lines. The host explicitly chooses whether to follow the returned remaining ranges; doing so is not a new ranking pass. The host owns question boundaries, stopping policy, and retained-context tracking.
Source reads accept UTF-8 with an optional BOM and normalize CRLF/CR to LF.
Only physical newlines define source coordinates: Unicode separators and formfeeds
inside source do not create extra line numbers. Explicit reads reject undecodable
or NUL-bearing input; content searches skip it with diagnostics. Path discovery
remains metadata-only. Root and nested .gitignore rules apply, including ignored
parent barriers; external directory junctions are not traversed.
Snapshot and API migration
Packs and receipts now use format 2. Rebuild old packs and reselect sources to create
new receipts; format-1 records are rejected, not silently certified against current
files. Receipt IDs address the request, source fingerprints, and rendered evidence.
Expansion rereads current source and reports expansion_changed_sources. These are
per-file snapshots, not an atomic working-tree snapshot or a full source archive.
Staged review uses a pinned index tree; --range A..B and A...B use B's pinned
source, including callers. External --diff uses working-tree source and cannot be
combined with those Git modes. JSON source_revision identifies the chosen source.
The unused controller/context-order modules and legacy Python candidate API were
removed. Use the supported selection/provider APIs; no compatibility shims remain.
The unused tree-sitter-python dependency and redundant benchmark sweep.py launcher
were also removed. Default MCP tools remain the same five.
CLI
pasr explain "how is the request rate limited" # run a selection, print the receipt
pasr trace enforce_per_user_request_quota src/ # a symbol's dependency closure
pasr trace HTTPAdapter src/ --callers # who calls it — impact analysis
pasr pack auth "session + login + token" src/auth/ # save a committable Context Pack
pasr review --staged src/ # touched defs + the callers they affect
pasr context --issue "$(cat issue.txt)" src/ \ # headless slice for CI / agents
--format text --metrics-file metrics.json
pasr report --price-per-mtok 3 # source-context estimates, not API billsThe source-reading commands above accept optional paths — with none, they scan
the whole workspace (.gitignore-aware). Pass directories or globs (src/,
lib/ "**/*.py") to scope them. doctor instead uses only its disposable fixture.
Receipt persistence to .pasr/receipts/<id>.{json,md} is best-effort (gitignored).
These records describe PASR selections, not a complete global agent audit. The
usage ledger at .pasr/ledger.jsonl supports source-context estimates, not measured
API savings or proven avoided round trips. Context Packs land in .pasr/packs/
(committable); review them for sensitive source before sharing. Load a named pack
with select_context(query="", advanced={"pack": "auth"}). For CI, see
docs/ci.md.
0.4.0 report compatibility
Version 0.4.0's report recomputes tokens_saved as signed
tokens_in - tokens_out, including for existing ledger rows. Negative values mean
context expansion; positive rows no longer hide larger outputs in other calls.
The ledger itself is not rewritten.
JSON reports include unmeasured_calls: rows without a valid nonnegative integer
for both token counters. An affected total or daily difference is null, rendered
as unknown; incomplete measurements produce neither a reduction percentage nor
a dollar estimate. A zero input total also has no defined reduction percentage.
Fully known counters and days remain reportable.
Consumers must accept nullable counters, differences, ratios and cost estimates,
and signed differences. Do not replace null with zero or clamp negative values;
check unmeasured_calls before interpreting totals. No ledger migration is needed.
--price-per-mtok must be finite and nonnegative; zero leaves the dollar estimate
disabled. JSON cost estimates retain sub-cent precision, and text displays eight
decimal places. These are hypothetical source-context differences, not provider
charges or measured savings. Coverage describes the recorded rows in the selected
date range, not missing writes or a complete agent session.
The reader refuses to generate a report from a damaged ledger: malformed or
unfinished JSON, non-object records, invalid UTF-8, or read failures produce CLI
exit code 2 and a diagnostic on stderr, with no text/JSON totals on stdout.
JSON/record errors identify the physical line without echoing its contents.
The whole file is checked before --since filtering, so a date filter cannot hide
corruption. Blank lines are ignored; Unicode separators inside valid query strings
remain part of that record. A genuinely absent ledger still means no recorded rows.
Reading never trims or repairs the file. If a writer is still appending, let it
finish before retrying; otherwise inspect a backup or restore a known-valid ledger.
These corrections change report semantics relative to 0.3.0.
Capability boundary
Use PASR to locate source and select inspectable context within a rendered budget. Computed spans and bounded static dependency traces need not contain the full mechanism. Global aggregation, lexical mismatch, dynamic bindings, and cross-file state transitions can require additional native searches or reads. Heuristic routing cannot guarantee detection of these gaps. Research results from other datasets do not establish answer-quality parity for this product.
Docs
examples/— five verbatim CLI transcripts against pinned reposeval/RESULTS.md— the 50-task evaluation, pre-registered (pip install ./eval→pasr-bench)docs/competitors-benchmark.md— latest real-component comparison, failed gates, and historical proxy resultsdocs/architecture.md— components and data flowdocs/roadmap.md— shipped and next
This productises the frozen researchv2 study (model-external context optimization); a
comparative write-up is in preparation.
Development
python -m venv .venv && . .venv/bin/activate # or .venv\Scripts\Activate.ps1
pip install -e ".[dev]"
pytest -qThe pure-logic core imports no torch / transformers (and no mcp SDK — that
loads only under pasr.mcp):
python -c "import pasr.pipeline, sys; assert not {'torch','transformers'} & set(sys.modules)"License
Available Tools
5 toolsfind_evidenceB
Locate query terms in workspace source, with ranked file:line hits and read ranges. include uses workspace-relative paths/globs: *.py is root-only; **/*.py is recursive.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| top_k | No | ||
| include | No | ||
| per_file | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does disclose scope (workspace source) and that results are ranked line-level hits with read ranges. It is silent on defaults (top_k=20, per_file=1), what an empty query does, and whether results are truncated or paginated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and result shape, followed by the one parameter detail that actually needs prose. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists (so return values needn't be restated), a 4-parameter, zero-annotation tool needs more than one parameter explained. The semantics of query, top_k, and per_file are absent from both schema and description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters, and the description only clarifies one of them (include), where its root-only vs recursive glob explanation is genuinely valuable. query, top_k, and per_file remain undocumented, so the description does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb+resource ("Locate query terms in workspace source") plus the shape of the result (ranked file:line hits, read ranges). It never names or contrasts with siblings like find_symbols, find_usages, find_files, or select_context, so an agent must infer the distinction on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The glob note implies the tool is a path-scoped text search, but there is no explicit when-to-use/when-not-to-use guidance and no mention of the alternatives (find_symbols for symbols, find_usages for references, find_files for filenames). The agent is left to guess which search tool fits.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_filesA
Rank workspace paths by query-term matches in their names. Does not search file contents. Optional include scopes paths; an empty query lists candidates.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| top_k | No | ||
| include | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does disclose real traits: matching is name-only (not content), ranking is by query-term matches, include scopes the path set, and an empty query enumerates candidates. It omits tie-breaking/ranking order, behavior on zero matches, and any auth or performance notes, so it is informative but not complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three terse sentences, zero filler, and the highest-value constraint ('Does not search file contents') is front-loaded immediately after the purpose statement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be described. The description covers the matching scope and the empty-query edge case, which is enough for an agent to invoke it correctly; the only real hole is the undocumented top_k parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains query (empty implies listing) and include ('scopes paths'), but top_k is never mentioned in either the schema or the description, leaving one of three parameters entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Rank workspace paths by query-term matches in their names,' and adds a scoping boundary ('Does not search file contents'). That boundary separates it from content-search siblings conceptually, but it never names find_evidence, find_symbols, find_usages or select_context, so the agent still infers which sibling to prefer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The exclusion 'Does not search file contents' and the note 'an empty query lists candidates' imply when the tool fits, but no alternative tool is named and no explicit when-to-use/when-not routing is given. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_symbolsA
Find exact or partial symbol definitions with file:line ranges (Python, JS/TS, Rust). include uses workspace-relative paths/globs: *.py is root-only; **/*.py is recursive.
| Name | Required | Description | Default |
|---|---|---|---|
| kinds | No | ||
| query | No | ||
| top_k | No | ||
| include | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses the return format ('file:line ranges'), the language scope, and glob traversal semantics, but says nothing about ranking, ordering, default query behavior, result caps, or what happens when the query is empty (the schema default).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the core capability ('find symbol definitions') is front-loaded before the glob detail. Every clause carries information the agent cannot get from structured fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The existence of an output schema excuses the description from explaining return values, and the glob semantics are a genuine addition. But for a 4-parameter tool with 0% schema coverage and no annotations, the omission of 'kinds' values, 'top_k' meaning, and explicit routing away from find_usages leaves the agent under-equipped.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters, so the description must compensate. It does real work for 'include' by explaining root-only vs recursive glob behavior, and 'exact or partial' hints at 'query' behavior, but 'kinds' and 'top_k' are left entirely opaque in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Find exact or partial symbol definitions', and adds output shape ('file:line ranges') plus supported languages (Python, JS/TS, Rust). It implies the distinction from find_usages (definitions vs references) but never names that sibling, so the differentiation is inferential rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'definitions' implicitly signals when to use this over find_usages, and the glob note gives practical usage guidance for scoping. However, there is no explicit when-to-use/when-not-to-use statement, no mention of the alternatives, and no prerequisites for the agent to anchor on.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_usagesB
Find lexical references to a symbol, with source lines and enclosing definitions when available. Returns at most top_k hits; read_lines suggests bounded source ranges.
| Name | Required | Description | Default |
|---|---|---|---|
| top_k | No | ||
| symbol | Yes | ||
| include | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose two useful traits: results are capped ('at most top_k hits') and source lines/enclosing definitions are only returned 'when available'. It adds a cross-tool hint (read_lines suggests bounded source ranges) but says nothing about permissions, the filtering behavior of `include`, or how hits are ordered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core action front-loaded and no filler. The second sentence efficiently bundles the result cap and the availability caveat, though the trailing read_lines reference slightly distracts from the tool's own contract.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the description does cover scope, cap, and availability. For a zero-annotation, zero-schema-coverage tool it still leaves gaps: the `include` parameter's meaning and the definition of a 'lexical' reference versus other search siblings are unresolved.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all three parameters. It explains top_k indirectly ('returns at most top_k hits') and `symbol` by implication, but `include` is never mentioned and its filtering semantics are left entirely undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('find lexical references to a symbol'), which is more precise than a generic 'search' and implicitly separates it from find_symbols (definitions) and find_evidence. It stops short of explicitly naming a sibling or drawing the boundary with them, so it is clear but not fully differentiating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the phrase 'find lexical references to a symbol', so an agent can infer the basic trigger. However, there is no explicit when-to-use, no exclusions, and no routing guidance relative to find_symbols, find_evidence, or select_context, leaving the choice between overlapping search tools to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
select_contextB
Select source spans from explicit files and/or include globs for a query. Context is capped at 1,500 tokens including file:line labels; metadata costs extra. Files accept path:start-end ranges. Range-only reads return ordered whole-line prefixes and continuation.files; whole-file or mixed scopes are ranked. outline returns a definition index. Each call is independent: retain prior context and compare source_fingerprint across continuations.
| Name | Required | Description | Default |
|---|---|---|---|
| files | No | ||
| query | Yes | ||
| include | No | ||
| outline | No | ||
| advanced | No | ||
| budget_tokens | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the 1,500-token cap (plus metadata cost), continuation semantics via continuation.files, the stateless nature of each call, and the need to compare source_fingerprint. This is substantive behavioral context an agent cannot get from structured fields alone, though reusability/caching caveats are not addressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first clause, and the following sentences are dense but each carries distinct operational information (token cap, range reads, outline, continuations). No filler, though the jargon-heavy phrasing demands re-reading in places.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is waived, and the description covers the cap, ranges, and continuation model for a fairly complex tool. Still missing is guidance on sibling selection and any meaning for query/advanced, so it is adequate but not fully complete for a 6-parameter tool with a nested object.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it partially does: files accepts path:start-end ranges, include takes globs, and outline returns a definition index (matching the 1500 default of budget_tokens). However, query, advanced (a nested free-form object), and the budget_tokens relationship remain unexplained, leaving half the parameters opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ("Select") and resource ("source spans") with the input sources (explicit files and/or include globs). It is reasonably distinguishable from siblings like find_files or find_evidence, though it never names an alternative to contrast against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance, nor any routing to the sibling tools (find_usages, find_evidence, find_symbols). The description explains mechanics (ranges, outline, continuations) but never says in what situation an agent should pick this tool over the others.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.3.0- First observed
find_evidence - First observed
find_files - First observed
find_symbols - First observed
find_usages - First observed
select_context
TDQS
Scored across 5 tools
find_evidence and select_context both locate source spans for a query, and find_usages/find_symbols overlap with find_evidence on symbol-related searches. Descriptions clarify intended use, but an agent may still misselect between grep-like and context-selection tools.
All tools use lowercase snake_case with a verb_noun structure (find_*, select_*). The single non-find verb (select) is a minor deviation given its distinct purpose.
Five tools is well-scoped for a read-only code navigation server. Each tool targets a distinct search or context-gathering operation without bloat.
The surface covers path search, symbol definitions, references, term search, and context selection, which is comprehensive for code navigation. Minor gaps include no explicit file-read or directory-listing tool, though select_context and find_files with empty query mitigate this.
Maintenance
Related MCP Connectors
Securely search and manage workspace context files for AI agents and teams.
Deterministic context layer for your codebase: change impact, blast radius, answers with receipts.
Project memory, semantic code search, and grounded agent context.
Codebase graphs, caller impact analysis, and recorded project context for AI coding agents.
Related MCP Servers
- FlicenseBqualityBmaintenanceEnables coding agents to search locally indexed repositories with hybrid semantic and lexical retrieval, returning exact source citations with file paths and line ranges.4-
- AlicenseBqualityAmaintenanceEnables coding agents to retrieve token-budgeted, source-hashed code context from explicitly registered local repositories via read-only MCP tools, reducing broad repository crawling while providing structure, symbols, and impact slices.10MIT
- AlicenseNot gradedqualityAmaintenanceProvides AI coding agents with token-budgeted, context-aware code retrieval from a persistent code graph, enabling efficient question answering without loading the entire repository.Apache 2.0
- AlicenseNot gradedqualityCmaintenanceEnables coding agents to retrieve a bounded context packet linking relevant code with prior session evidence, decisions, failures, and fixes from local Codex and Claude Code history.Apache 2.0