library-enrichment
Provides library-evidence and dependency research for Python packages, including exact release resolution, API/doc/source evidence search, symbol inspection, release comparison, and isolated usage verification.
Provides library-evidence and dependency research for Rust crates, including exact release resolution, API/doc/source evidence search, symbol inspection, release comparison, and isolated usage verification.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@library-enrichmentresolve serde 1.0.203 and search its derive macro evidence"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
library-enrichment
A local-first library-evidence service for Rust and Python, exposed over MCP.
It answers the question a package index cannot: given a design objective and a concrete dependency environment, what relevant library capabilities exist, what evidence supports their use, what configuration do they require, and what remains unverified?
The service produces evidence. The calling agent produces judgments. It complements Context7 — Context7 explains concepts, this service establishes exact-release facts about the dependencies a project actually resolved.
Status
Phase 0 of 6. See STATUS.md for the current gate tally and
docs/reports/acceptance.md for per-gate results. This repository is under
active construction; the governing specification is complete and frozen, the implementation
is not.
Related MCP server: voyager
What it does
Nine MCP tools over a shared evidence core:
Tool | Purpose |
| Establish exact release identity and environment before research |
| Discover unfamiliar capabilities without knowing symbol names |
| Search a bounded set of API, docs, example, source, or release evidence |
| Characterize a candidate and its deployment requirements |
| Discover additions, removals, and non-API changes |
| Test a proposed invocation in an isolated capsule |
| Retrieve large result sections without flooding context |
| Observe or cancel a long-running operation |
| Inspect readiness and capabilities |
Evidence carries provenance: every material fact traces to an exact artifact, a producer run, and a recorded environment. Declared availability, configured availability, type-check success, and runtime behavior are separate states — the service never collapses them.
Architecture
A Rust daemon owns identity, resolution, fetching, normalization, evidence storage, querying, job state, policy, and publication. Evidence lands in immutable Arrow/Parquet snapshots queried through DataFusion. A thin FastMCP 4 adapter exposes the tools; a separate worker runs Griffe extraction and isolated runtime probes.
coding agent ──┬── Context7 MCP ─────────► concepts, documentation examples
│
├── library-research skill ► routing, evidence policy, synthesis
│
└── FastMCP 4 stdio adapter
│ typed, bounded local RPC
▼
library-enrichmentd ──► fetchers, Rust/Python API producers,
│ LSP sessions, sandboxed verification
▼
immutable evidence snapshots (Arrow/Parquet + DataFusion)Service code, tool environments, extraction workspaces, caches, and outputs live outside every working repository. The repository under study is never a subprocess working directory, an extraction destination, or an install target.
Documentation
Path | Contents |
The governing specification — frozen, not edited | |
Compatibility matrix and design notes | |
Decision records, including every deviation from the blueprint | |
Install, configure, and run | |
Frozen response-envelope schema and fixtures | |
The 46 acceptance gates | |
The companion agent skill | |
Instructions for agents working in this repository |
Development
just --list # every recipe
just doctor # verify the toolchain and required producers are present
just ci # gates for completed phases, plus the next oneRequires a pinned stable Rust toolchain (rust-toolchain.toml), uv, and just.
Build and runtime verification profiles additionally require a working container runtime
or bwrap.
License
Dual-licensed under either MIT or Apache-2.0, at your option.
Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in this work by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.
Available Tools
9 toolscompare_releasesCompare ReleasesCIdempotent
Discover additions/removals and non-API changes.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| cursor | No | ||
| scopes | No | ||
| ecosystem | No | ||
| max_bytes | No | ||
| max_items | No | ||
| to_version | No | ||
| from_version | No | ||
| after_context_id | No | ||
| after_snapshot_id | No | ||
| before_context_id | No | ||
| alternative_cursor | No | ||
| before_snapshot_id | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| job | Yes | |
| data | Yes | |
| error | Yes | |
| status | Yes | |
| summary | Yes | |
| coverage | Yes | |
| delivery | Yes | Delivery changes representation, never the original research status or coverage. |
| evidence | Yes | |
| artifacts | Yes | |
| freshness | Yes | |
| context_id | Yes | |
| request_id | Yes | |
| snapshot_id | Yes | |
| schema_version | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds no behavioral detail beyond the annotations. It does not mention pagination, cursor usage, response format, or any constraints like requiring from_version and to_version. The annotations (openWorldHint, idempotentHint) provide some context but the description does not build on them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence with no filler. It is appropriately concise and easy to parse, though it lacks the substance needed for a tool with this many parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (13 parameters, no required fields, multiple optional filters), the description is far too sparse. It does not explain what 'additions/removals' refers to, how scopes or ecosystem affect results, or what the output looks like. The output schema exists but does not help with parameter usage. The tool is under-specified for an agent to use effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description provides no explanation for any of the 13 parameters. It does not clarify the meaning of scopes, ecosystem, max_bytes, or the various ID fields. The description completely fails to compensate for the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (discover) and a clear object (additions/removals and non-API changes). It clearly implies comparing releases, and this is distinct from sibling tools like library_overview or search_evidence. However, it doesn't explicitly mention comparing two versions or mention from/to version parameters, so it is not maximally specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus any alternative. There is no mention of appropriate scenarios, prerequisites, or exclusions. The agent is left to infer usage from the name and description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_symbolInspect SymbolA
Read retained symbol evidence; explicit execution options can run isolated semantic or runtime inspection.
| Name | Required | Description | Default |
|---|---|---|---|
| execution | No | Read retained evidence or explicitly select an execution profile. | |
| max_bytes | No | Inline byte budget; the server caps it. | |
| selection | No | ||
| context_id | Yes | From `resolve_library`. | |
| snapshot_id | No | Read a specific snapshot; the context's current one otherwise. | |
| symbol_path | Yes | Fully qualified symbol path. | |
| definition_id | No | Select a definition from candidates when the public path is ambiguous. |
Output Schema
| Name | Required | Description |
|---|---|---|
| job | Yes | |
| data | Yes | |
| error | Yes | |
| status | Yes | |
| summary | Yes | |
| coverage | Yes | |
| delivery | Yes | Delivery changes representation, never the original research status or coverage. |
| evidence | Yes | |
| artifacts | Yes | |
| freshness | Yes | |
| context_id | Yes | |
| request_id | Yes | |
| snapshot_id | Yes | |
| schema_version | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds useful behavior beyond the annotations: retained reads are safe, but 'explicit execution options can run isolated semantic or runtime inspection,' which explains why readOnlyHint is false and warns that execution is possible. It does not detail side effects or the runtime capsule behavior, but the schema covers those details, so the description still earns credit for shaping expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight clauses with no filler. It front-loads the main read action and then adds the important execution fork, making it efficient and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a complex 7-parameter schema with detailed per-parameter descriptions and an output schema, so the description does not need to restate everything. However, it omits high-level context such as the prerequisite relationship to resolve_library and when to choose execution profiles, leaving some gaps that an agent must infer or discover from the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high (86%), so the schema carries most parameter semantics. The description's phrase about execution options roughly maps to the execution parameter, but it does not add meaning beyond what input schema descriptions already state. A baseline 3 is appropriate here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Read retained symbol evidence,' which distinguishes symbol inspection from sibling tools like read_artifact or search_evidence. It also notes a secondary behavior (explicit execution for semantic/runtime inspection), giving a clear functional identity. It stops short of a 5 because it does not explicitly contrast with sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives such as search_evidence, resolve_library, or read_artifact. The phrase 'retained symbol evidence' implies a use case, but the description does not state prerequisites, exclusions, or when the execution options should be selected.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
job_controlJob ControlA
Observe or cancel an explicitly submitted long-running operation.
| Name | Required | Description | Default |
|---|---|---|---|
| action | No | `cancel` drops your interest; shared work may continue. | status |
| job_id | Yes | From a `pending` result. | |
| max_bytes | No | ||
| wait_seconds | No | Bounded wait for `wait`. | |
| interest_token | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| job | Yes | |
| data | Yes | |
| error | Yes | |
| status | Yes | |
| summary | Yes | |
| coverage | Yes | |
| delivery | Yes | Delivery changes representation, never the original research status or coverage. |
| evidence | Yes | |
| artifacts | Yes | |
| freshness | Yes | |
| context_id | Yes | |
| request_id | Yes | |
| snapshot_id | Yes | |
| schema_version | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is consistent with annotations: 'observe or cancel' aligns with readOnlyHint=false (cancel is a mutation) and does not contradict destructiveHint=false. However, it adds little behavioral context beyond the annotations themselves. The schema's 'cancel drops your interest; shared work may continue' is valuable context but lives in the schema, not the description. With openWorldHint=true signaling potential external effects, the description could have disclosed more about side effects but doesn't.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler words. The core action (observe/cancel) leads and the resource scope (explicitly submitted long-running operation) follows. Efficient, though so terse it borders on under-specification rather than disciplined conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists so return values need not be described, and annotations carry the safety profile. But for a 5-parameter tool with three distinct actions (status/wait/cancel), the description omits any guidance on which action to select and when. The presence of wait_seconds and interest_token suggests nontrivial interaction patterns that are undocumented in the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
At 60% schema coverage, two parameters (max_bytes, interest_token) lack schema descriptions and the description compensates for none of them. The description contains zero parameter information. An agent cannot infer what max_bytes or interest_token mean from either the schema or the description — a meaningful gap the description fails to bridge.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Observe or cancel an explicitly submitted long-running operation' uses a specific verb pair and a precise resource type. It cleanly distinguishes itself from all eight siblings (service_status, library_overview, search_evidence, etc.), none of which deal with job lifecycle management — an agent can select this tool unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'explicitly submitted' implies this tool is for jobs that were returned as pending, and the schema reinforces this with job_id described as 'From a pending result.' However, the description itself gives no explicit when-to-use guidance, no exclusions, and no named alternatives — the usage context is only implied, not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
library_overviewLibrary OverviewBRead-onlyIdempotent
Discover unfamiliar capabilities without knowing symbol names.
| Name | Required | Description | Default |
|---|---|---|---|
| area | No | Narrow the map to one module or feature area. | |
| discovery | No | Independent feature, README, release-note and example pages. Omit for bounded previews; use an empty list for namespace navigation alone. | |
| max_bytes | No | Inline byte budget; the server caps it. | |
| max_items | No | Child entries per namespace; the server caps it. | |
| context_id | Yes | From `resolve_library`. | |
| snapshot_id | No | Read a specific snapshot; the context's current one otherwise. |
Output Schema
| Name | Required | Description |
|---|---|---|
| job | Yes | |
| data | Yes | |
| error | Yes | |
| status | Yes | |
| summary | Yes | |
| coverage | Yes | |
| delivery | Yes | Delivery changes representation, never the original research status or coverage. |
| evidence | Yes | |
| artifacts | Yes | |
| freshness | Yes | |
| context_id | Yes | |
| request_id | Yes | |
| snapshot_id | Yes | |
| schema_version | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, and the description does not contradict them. The description adds no behavioral context about pagination, namespace navigation, snapshot handling, or server caps; an agent must read the schema to learn these traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one compact sentence with no filler or repetition. It is highly scannable, although it conveys more of a tagline than a detailed specification.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema and annotations, this tool has six parameters including a nested discovery structure, and the description gives no guidance on when to set discovery, how namespaces are navigated, or how snapshots affect results. An agent would need to infer most operational behavior from the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with each parameter including constraints and defaults. The description adds no parameter-level meaning, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource pair ('discover unfamiliar capabilities') and implies differentiation from symbol-oriented siblings like inspect_symbol. However, it does not state the concrete outputs (feature pages, docs, examples) or mention the discovery parameter, leaving the purpose somewhat high-level.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without knowing symbol names' implies use for exploration when symbol names are unknown, which distinguishes it from inspect_symbol. But no explicit when-to-use or when-not-to-use guidance is given, and no sibling alternatives are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_artifactRead ArtifactARead-onlyIdempotent
Retrieve large result sections without flooding context.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | Continue a previous read. | |
| section | No | Read a typed result section or a Markdown heading. | |
| max_bytes | No | Bytes per slice; the server caps it. | |
| artifact_id | Yes | From a result's `artifacts`. |
Output Schema
| Name | Required | Description |
|---|---|---|
| job | Yes | |
| data | Yes | |
| error | Yes | |
| status | Yes | |
| summary | Yes | |
| coverage | Yes | |
| delivery | Yes | Delivery changes representation, never the original research status or coverage. |
| evidence | Yes | |
| artifacts | Yes | |
| freshness | Yes | |
| context_id | Yes | |
| request_id | Yes | |
| snapshot_id | Yes | |
| schema_version | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is established. The description adds the 'without flooding context' behavior, but it does not describe failure behavior, cursor limits, or what happens when the server caps bytes; the schema carries that weight.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler. 'Retrieve large result sections without flooding context' conveys purpose and usage in one line.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the rich schema, output schema presence, and read-only/idempotent annotations, the description is sufficient for an agent to invoke the tool correctly. It does not explain the distinction between markdown-heading sections and typed result sections, but the schema already does that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with clear explanations for artifact_id, cursor, section, and max_bytes. The description itself adds no parameter-specific meaning, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description says exactly what the tool does—retrieve large result sections—and adds the outcome ('without flooding context'). It is specific about the verb and resource, though it does not explicitly contrast itself with sibling tools like inspect_symbol or search_evidence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Retrieve large result sections without flooding context' supplies a clear condition for when to use this tool: when a result is large enough that context needs careful pagination. It names no exclusions or alternatives, so it is useful but not fully explicit about when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_libraryResolve LibraryCIdempotent
Establish exact identity and environment before research.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Research mode. Defaults to `project` with a version and `upstream` without. | |
| name | Yes | Crate or distribution name. | |
| extras | No | ||
| target | No | Your project's target triple or platform, if known. | |
| version | No | Exact version. Omit only for an explicit upstream question. | |
| features | No | Rust features your project enables; use extras for Python. | |
| revision | No | ||
| ecosystem | Yes | Which package ecosystem. | |
| freshness | No | `cache_ok` retains exact-version evidence indefinitely; latest selection has a TTL. `revalidate` always consults the registry; `offline` never opens a socket. | cache_ok |
| repository | No | ||
| allow_yanked | No | ||
| package_subdir | No | ||
| python_version | No | ||
| allow_prerelease | No | ||
| default_features | No | Whether your project enables default features, if known. | |
| allow_local_build | No | Accept a local rustdoc build when docs.rs JSON is missing or in an unreadable format, or its observed build differs from requested features/target. Compiles the crate on a dated nightly in an isolated capsule and can take minutes. Requires the operator-enabled `build` profile; asking never grants it. |
Output Schema
| Name | Required | Description |
|---|---|---|
| job | Yes | |
| data | Yes | |
| error | Yes | |
| status | Yes | |
| summary | Yes | |
| coverage | Yes | |
| delivery | Yes | Delivery changes representation, never the original research status or coverage. |
| evidence | Yes | |
| artifacts | Yes | |
| freshness | Yes | |
| context_id | Yes | |
| request_id | Yes | |
| snapshot_id | Yes | |
| schema_version | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare idempotentHint=true and destructiveHint=false, so there is no contradiction here. However, the description adds no behavioral context beyond the outcome: it does not disclose network usage, caching, writes, build-time cost, or failure modes, which matters given readOnlyHint=false and openWorldHint=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler and is easy to parse. It may be too spare, but its brevity is still a strength.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter tool without readOnlyHint and with many research-oriented siblings, a one-sentence description is inadequate. It does not explain what 'environment' includes, when this tool should be called relative to others, or what inputs drive the resolution, despite an output schema existing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 56%, and the description provides no parameter-level guidance. Several parameters such as revision, repository, allow_yanked, package_subdir, and python_version are not described in the schema or the description, so the description does not compensate for the gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The phrase 'Establish exact identity and environment before research' names a concrete outcome (resolved package identity/environment) and frames it as a pre-research step. It is not a tautology and conveys a clear job for the agent, though it does not explicitly differentiate it from siblings like library_overview or search_evidence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Before research' implies this tool should be used first, but the description never states when to choose it over library_overview, search_evidence, or compare_releases, nor any exclusions. Usage context is only implied, not explicitly specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_evidenceSearch EvidenceBRead-onlyIdempotent
Search a bounded set of API, docs, examples, source, or release evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| area | No | Restrict to this namespace subtree. | |
| kinds | No | Restrict to evidence kinds; all but `source` when omitted. | |
| query | Yes | What to look for. | |
| cursor | No | Continue a previous search. Same query, filters and snapshot. | |
| max_bytes | No | Inline byte budget; the server caps it. | |
| max_items | No | Page size; the server caps it. | |
| context_id | Yes | From `resolve_library`. | |
| snapshot_id | No | Read a specific snapshot; the context's current one otherwise. |
Output Schema
| Name | Required | Description |
|---|---|---|
| job | Yes | |
| data | Yes | |
| error | Yes | |
| status | Yes | |
| summary | Yes | |
| coverage | Yes | |
| delivery | Yes | Delivery changes representation, never the original research status or coverage. |
| evidence | Yes | |
| artifacts | Yes | |
| freshness | Yes | |
| context_id | Yes | |
| request_id | Yes | |
| snapshot_id | Yes | |
| schema_version | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnlyHint, idempotentHint, and closed-world behavior, so the safety profile is established. The description adds the evidence categories and the 'bounded set' limitation, which is useful, but it does not mention pagination, snapshot semantics, or default filtering behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It clearly starts with the action 'Search' and immediately conveys the resource scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with eight parameters, cursor-based pagination, snapshot semantics, and default kind filtering, the one-sentence description is too thin. It does not explain what bounds the set, how context_id relates to resolve_library, or the overall search model, so the agent must rely on the schema to understand operational details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all eight parameters, including query, context_id, kinds, cursor, max_bytes, max_items, and snapshot_id. The description adds no parameter-level meaning, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'search' and identifies a concrete resource: 'a bounded set of API, docs, examples, source, or release evidence.' This distinguishes it from siblings like inspect_symbol or read_artifact, though it does not explicitly contrast with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus siblings such as verify_usage, inspect_symbol, or compare_releases. It states only what the tool does, leaving the agent to infer appropriate usage from the name and parameters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
service_statusService StatusBRead-onlyIdempotent
Inspect readiness and capabilities, without indexing.
| Name | Required | Description | Default |
|---|---|---|---|
| component | No | Restrict the report to one producer or component. |
Output Schema
| Name | Required | Description |
|---|---|---|
| job | Yes | |
| data | Yes | |
| error | Yes | |
| status | Yes | |
| summary | Yes | |
| coverage | Yes | |
| delivery | Yes | Delivery changes representation, never the original research status or coverage. |
| evidence | Yes | |
| artifacts | Yes | |
| freshness | Yes | |
| context_id | Yes | |
| request_id | Yes | |
| snapshot_id | Yes | |
| schema_version | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, covering the safety profile. The description adds 'without indexing' which hints at a non-mutating, non-heavy operation, but it doesn't elaborate on what readiness/capabilities entail or any side effects. Given the annotations carry most of the burden, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It is appropriately concise for a simple tool, though it could be slightly more informative. No waste, but it doesn't earn a 5 due to its brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the presence of an output schema (which covers return values), high schema coverage for the one parameter, and annotations covering safety and idempotency, the description is nearly complete. It lacks elaboration on what 'readiness and capabilities' means, but that is likely detailed in the output schema. This is sufficient for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description for the single parameter 'component' has 100% coverage, explaining it restricts the report. The tool description adds no additional meaning about parameters, so with full schema coverage the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Inspect') and resource ('readiness and capabilities'), which is clear. It also adds 'without indexing' as a distinguishing qualifier, but it does not explicitly differentiate from the sibling tools like inspect_symbol or library_overview. It's not a tautology and gives a concrete sense of the operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus its siblings. The description simply states what it does but never mentions alternatives, prerequisites, or context that would help an agent choose it over other inspection tools. This is a clear gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_usageVerify UsageA
Test a proposed invocation or composition in isolation.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | What to establish. These prove different things. | typecheck |
| profile | No | build | |
| snippet | Yes | The code to verify. | |
| max_bytes | No | ||
| context_id | Yes | From `resolve_library`. | |
| snapshot_id | No | ||
| test_intent | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| job | Yes | |
| data | Yes | |
| error | Yes | |
| status | Yes | |
| summary | Yes | |
| coverage | Yes | |
| delivery | Yes | Delivery changes representation, never the original research status or coverage. |
| evidence | Yes | |
| artifacts | Yes | |
| freshness | Yes | |
| context_id | Yes | |
| request_id | Yes | |
| snapshot_id | Yes | |
| schema_version | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal that the tool is not read-only and not idempotent; the description adds that testing happens 'in isolation,' which is useful but not detailed. It does not disclose whether code is actually executed, what side effects may occur, or whether resource limits apply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence delivers the core purpose with zero filler. It is concise at the cost of depth, but that depth is penalized in other dimensions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 7 parameters, 2 enums, and an output schema, a one-sentence description is insufficient to guide safe invocation. It omits the workflow (e.g., context_id from resolve_library appears only in the schema), the purpose of modes/profiles, and any caveats about executing snippets.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 43%, with max_bytes, profile, snapshot_id, and test_intent left undocumented. The description adds no parameter-level meaning and therefore does not compensate for these gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Test') and a clear resource ('a proposed invocation or composition in isolation'). This distinguishes it from sibling tools like service_status, resolve_library, and inspect_symbol, which serve different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a pre-flight testing use case but does not explicitly state when to use verify_usage instead of siblings or when to avoid it. No exclusions or alternative conditions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.0.0- First observed
compare_releases - First observed
inspect_symbol - First observed
job_control - First observed
library_overview - First observed
read_artifact - First observed
resolve_library - First observed
search_evidence - First observed
service_status - First observed
verify_usage
TDQS
Scored across 9 tools
Each tool maps to a distinct phase in the research workflow: status, identity, overview, evidence search, symbol inspection, release comparison, usage verification, artifact retrieval, and job control. Even the two preflight tools are clearly separated by readiness vs. identity.
The majority use verb_noun (search_evidence, inspect_symbol, compare_releases, verify_usage, read_artifact), but service_status, library_overview, and job_control are noun_noun, creating an inconsistent pattern. Names remain readable, so not chaotic.
Nine tools is a well-scoped count for a specialized library-research server; each tool covers a distinct step without redundancy or bloat.
The set covers the full discovery, inspection, comparison, verification, and retrieval lifecycle, with job control for long operations. The only minor gap is the absence of an explicit write-back/enrichment-submission tool, though the server appears read-oriented by design.
Maintenance
Related MCP Connectors
Machine-native research commons for agent evidence, discovery, rooms, and bounded research quests.
Search GitHub, npm, PyPI, StackOverflow, ArXiv from one MCP — built for coding agents.
Cross-agent artifact workspace with provenance across Claude Code, Codex, Cursor, LangGraph.
Versioned documentation registry and semantic search for AI tools and coding assistants.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceProvides up-to-date, version-specific documentation and code examples for software libraries directly to LLMs, enabling resolution of library identifiers and retrieval of relevant documentation with code snippets.-
- AlicenseNot gradedqualityBmaintenanceAn MCP server that enables coding agents to safely verify npm/PyPI packages, retrieve cited briefs, discover GitHub repos, and fetch canonical docs, with fail-closed OSV vulnerability checks and injection-hardened output.166 npmMIT

mentu-navigator-mcpofficial
AlicenseNot gradedqualityCmaintenanceProvides read-only, provenance-first repository navigation for agents and humans, with ranked lexical retrieval, exact query, document handles, symbol context, and change impact analysis.75 npmApache 2.0- AlicenseNot gradedqualityBmaintenanceEnables AI coding agents to search and retrieve cited line-level context from versioned knowledge bundles, as well as list workspaces and publish or pull immutable Markdown bundles.Apache 2.0