PHION Agent Trust Infrastructure
Server Details
123 agent services with free discovery, intent, x402 terms and signed verification.
- Status
- Healthy
- Uptime
- 88.0% over 23 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
TDQS
Scored across 135 tools
Many tools have overlapping purposes, especially around payment preflight (payment_preflight, payment_capability_preflight, x402_v2_router, payment_route_selector) and transaction recovery (transaction_recovery, transaction_recovery_v2). Dozens of tools share the identical vague description 'deterministic signed assessment over supplied input', making boundaries unclear. An agent would frequently misselect among similar-sounding options.
Nearly all tool names use consistent snake_case, with clear noun/verb phrases. Minor deviations exist (single-word names like preflight, verify, attest; some with phion_ prefix), but the overall pattern is predictable and readable.
135 tools is an extreme mismatch for any coherent server purpose. This far exceeds the 50+ threshold for a score of 1, and the sheer volume creates redundancy and maintenance burden rather than focused capability.
The surface covers a wide range of trust-infrastructure domains: payment routing, identity, reputation, enrichment, RWA compliance, transaction assurance, and discovery. While some lifecycle operations are implicit or missing, no obvious critical gap would block an agent from core trust and payment workflows.
Available Tools
135 toolsa2a_transaction_bridgeA2A Transaction BridgeARead-onlyIdempotentInspect
Bind an A2A task to a transaction without silently transferring authority.; 0.004 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, and non-destructive behavior, and the description adds genuinely useful context beyond those: the 0.004 USDC cost, determinism, evidence-boundedness, and signed output. It does not detail what the fee pays for or how the 'evidence-bounded' guarantee manifests, but it complements the annotation profile rather than merely repeating it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short fragments each carry distinct information: the core action, the cost, and the tool's guarantees. The description is front-loaded and every phrase earns its place, though the mixed semicolon/period punctuation is slightly awkward.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with an open input schema and several close siblings, the description leaves meaningful gaps: what '0.004 USDC' refers to (fee vs bound amount), what the required input properties are, what 'evidence-bounded' means operationally, and when to prefer this over mcp_a2a_task_bridge. The existence of an output schema covers return values, but the input and selection contract is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero named parameters, the schema provides no constraints, and a 0-param baseline of 4 would normally apply. But the schema is an open object with minProperties 1, so the agent must guess which properties to supply. The description hints at relevant entities (A2A task, transaction, USDC, authority) but never maps them to concrete input fields, leaving the input contract underspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Bind an A2A task to a transaction') plus a key constraint ('without silently transferring authority'), which gives the agent a clear action and a distinguishing negative. It does not, however, explain the 'A2A' jargon or differentiate itself from the close sibling mcp_a2a_task_bridge, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without silently transferring authority' implies a usage condition: use this when you want to associate a task with a transaction but not hand over authority. However, the description never names alternatives or states when not to use the tool, and the sibling list contains several overlapping bridges (mcp_a2a_task_bridge, ap2_mandate_bridge) that are left unaddressed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
abandoned_tool_detectorAbandoned Tool DetectorCRead-onlyIdempotentInspect
Require stale success plus multiple independent failed probes before classifying likely abandonment; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| tool | Yes | ||
| policy | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, covering the safety profile. The description adds two behavioral facts beyond annotations: the conservative gating logic (stale success plus multiple independent failed probes) and a per-call cost of 0.003 USDC. It lacks details on auth, rate limits, or what happens when conditions are not met.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single semicolon-joined sentence with no wasted words, but it front-loads the condition rather than the action. Its terse, under-specified style leaves important gaps, so it is concise but not optimally structured for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a detector that takes a nested policy object, has no output schema, and has 0% schema description coverage, the description omits what the policy object contains, what the output looks like, and how the classification is returned. It provides only cost and a high-level condition, which is insufficient for confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and there are two required nested parameters. The description partially maps conditions to the 'tool' object's fields ('stale success' → last_success_at, 'multiple independent failed probes' → independent_probe_count/reachable), but it says nothing about the required 'policy' object, leaving half the parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description indicates the tool classifies likely abandonment but does so indirectly through a condition ('Require stale success plus multiple independent failed probes before classifying...') rather than a direct verb like 'detects' or 'classifies'. It does not differentiate from siblings such as tool_capability_drift_monitor. The pricing detail ('0.003 USDC') is not purpose-related.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance is given. The conditional phrasing describes internal logic, not invocation context, and no alternative tools are named. The agent must infer that this is for abandonment classification.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_approval_relayAgent Approval RelayCRead-onlyIdempotentInspect
Agent Approval Relay; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds useful behavior beyond the annotations: the result is deterministic and signed, and it operates only over the supplied input. It does not contradict the readOnly/idempotent hints, though it omits details about signature format, validation, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely short, but it begins by repeating the tool name and title. The remaining phrase adds some value, but the entry is under-specified rather than efficiently informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite simple schema and safe annotations, the tool cannot be invoked confidently because the meaning of 'approval' and the expected input shape are never explained. An agent would not know whether to pass a boolean, an object, or some other structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high, so the baseline is 3; the description does not need to re-list parameter details. That said, the required 'approval' parameter is completely opaque, and the description's 'supplied input' does not clarify what form the approval should take.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description says it produces a 'deterministic signed assessment over supplied input,' which identifies a specific kind of output and a clear scope. However, it is vague about what 'approval relay' actually does and does not distinguish it from related tools like attest, verify, or signed_result_comparator.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus any of its many siblings. The description neither states the intended use case nor names alternatives, so an agent cannot decide when this is the right choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_behavior_fingerprintAgent Behavior FingerprintARead-onlyIdempotentInspect
Create a privacy-minimized behavioral commitment without retaining payloads or raw identifiers.; 0.003 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context: it does not retain payloads or raw identifiers, is deterministic, evidence-bounded, and signed, and includes a cost (0.003 USDC). This goes beyond annotations and helps an agent understand side effects and guarantees.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no waste. It front-loads the core purpose, then adds the cost and key properties. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the open schema, no required parameters, and rich annotations, the description is fairly complete. It explains the behavioral commitment, privacy properties, cost, and signing. It doesn't describe the output schema, but an output schema exists, so that's not required. A small gap is the lack of explicit guidance on what inputs are needed, but the open schema makes that less critical.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema is open (additionalProperties: true, minProperties: 1) with no named parameters, so schema coverage is effectively 100% but provides no semantic detail. The description compensates by explaining the purpose and constraints (privacy-minimized, no raw identifiers, deterministic, evidence-bounded, signed), which gives an agent enough context to know what kind of input is expected. However, it doesn't specify exact required fields, so a 4 is appropriate rather than 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Create') and a clear resource ('a privacy-minimized behavioral commitment'), and distinguishes it from siblings by emphasizing privacy-minimization and evidence-boundedness. It doesn't explicitly name a sibling alternative, but the phrase 'without retaining payloads or raw identifiers' clarifies its unique scope among the many evidence/receipt tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: when a privacy-minimized, deterministic, signed behavioral commitment is needed. It does not explicitly state when not to use it or name alternatives, but the context of sibling tools (e.g., agent_reputation_evidence, attest) suggests a niche. The guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_budget_guardCInspect
Enforce per-call, task, session and period budgets; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the cost (0.002 USDC) and budget scopes, but says nothing about what happens when a budget is exceeded, whether the operation blocks or logs, required permissions, idempotency, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is front-loaded with the core action and scope, with zero wasted words. It is appropriately sized but extremely terse, bordering on under-specification rather than being structurally deficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema, the definition is incomplete for a guard tool with no annotations: it omits critical when-to-use context, behavioral outcomes on enforcement, and any differentiation from siblings. An agent has enough to know the broad purpose but not enough to call it correctly in context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is reported at 100%, so the structured schema is expected to document parameters, setting the baseline at 3. The description adds no mapping or meaning for the required request, usage, and policy parameters beyond the general budget categories already implied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Enforce') and resource ('budgets'), and enumerates per-call, task, session and period scopes, making the core purpose clear. However, it does not differentiate itself from similar sibling guards like subscription_spend_guard or mandate_reserve, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives, nor any stated exclusions or prerequisites. The budget enforcement context implies a use case, but an agent must infer when to call it relative to the many other guard and payment tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_economic_graphAgent Economic GraphARead-onlyIdempotentInspect
Build privacy-bounded economic relationships only from supplied nodes and edges.; 0.005 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnly, idempotent, and non-destructive hints. The description adds meaningful behavioral context beyond those: a fixed fee of 0.005 USDC, determinism, evidence-boundedness, signing, and the restriction to only supplied data. No contradiction with annotations is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief and front-loaded with the primary purpose alert. The three clauses each add distinct information: what it builds, what it costs, and how it behaves. The punctuation/format is slightly awkward, but there is no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no defined schema parameters, an output schema, and rich annotations, the description covers the essential expectations: construction, privacy boundaries, cost, determinism, evidence, and signing. It lacks usage-direction details, but those belong to the usage dimension. Overall, an agent likely has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has no explicit parameters, so the 0-param baseline of 4 applies. The description compensates somewhat by indicating the input should consist of supplied nodes and edges, even though it does not define exact property names. Since the schema itself is an open object, the description adds at least directional meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Build') and resource ('privacy-bounded economic relationships') and constrains scope to 'supplied nodes and edges'. It is distinct from more generic graph or relationship tools, though it does not explicitly name sibling alternatives. The added attributes 'Deterministic, evidence-bounded and signed' further clarify the output's nature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives. The phrase 'only from supplied nodes and edges' implies a privacy/scope constraint but does not state conditions, exclusions, or which sibling tools are more appropriate in other scenarios. An agent would have to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_memory_provenanceAgent Memory ProvenanceCRead-onlyIdempotentInspect
Agent Memory Provenance; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive behavior, so the safety profile is covered. The description adds 'deterministic signed,' which hints at a reproducibility and output-signing property, but it does not clarify what 'signed' means here or what side effects, if any, the assessment has on the supplied input.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded, with no wasted phrasing. However, it is terse to the point of under-specification, leaving important semantic and selection information absent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the large, semantically overlapping sibling list and unnamed required parameters, the description is incomplete for reliable tool selection and invocation. The presence of an output schema helps, but the description still fails to convey the intended use case, input semantics, or differentiation criteria.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The required parameters 'memory' and 'source' are given no semantic explanation in either the schema or the description. The phrase 'supplied input' is generic and does not clarify which parameter maps to which concept, so an agent cannot confidently determine what values to provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description says the tool produces a 'deterministic signed assessment over supplied input,' which conveys the basic operation and resource. However, it does not explain what 'agent memory provenance' means, what question is being answered, or how this differs from many similarly named provenance/attestation siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus alternatives. The sibling list contains numerous provenance, evidence, and attestation tools, but the description provides no conditions, prerequisites, or exclusions to help an agent choose among them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_rate_limit_negotiatorAgent Rate Limit NegotiatorCRead-onlyIdempotentInspect
Agent Rate Limit Negotiator; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide rich safety context: readOnlyHint=true, idempotentHint=true, destructiveHint=false. The description adds only 'deterministic' (which overlaps with idempotentHint) and 'signed,' which is undefined. There is no contradiction with annotations, but the added behavioral insight is minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short, but brevity here is under-specification rather than useful compression. It repeats the tool name and adds a vague clause without front-loading any concrete, actionable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With two required opaque parameters, no property definitions, and no explanation of the negotiation semantics, the description is insufficient for confident invocation. The output schema exists, but it does not compensate for missing input semantics and absent usage context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema requires 'demand' and 'offer' but provides no property definitions, and the description's phrase 'supplied input' adds no meaning to either parameter. An agent is left to guess what shape demand and offer should take, what they represent, and whether additional properties are meaningful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description moves slightly beyond a pure tautology by saying the tool produces a 'deterministic signed assessment over supplied input,' but it never states a concrete verb or resource. It does not clarify whether the tool negotiates rate limits, assesses a demand/offer pair, or returns a signed verdict, and it does not distinguish it from sibling negotiation or assessment tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus alternatives such as capability_negotiation_preflight or agent_budget_guard. No conditions, exclusions, or examples are given; the only implied use case comes from the tool name itself, not the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_reputation_evidenceCInspect
Evidence-bounded agent reputation assessment with confidence, coverage, provenance limitations and a signed receipt; 0.005 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| signals | No | ||
| subject | Yes | ||
| evidence | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose two useful behavioral traits: the assessment is bounded by supplied evidence, and the call costs 0.005 USDC (a paid operation). It does not disclose what happens when evidence is insufficient, what the signed receipt guarantees, or any permission/latency characteristics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with no filler, and the core purpose plus the output contents and price are front-loaded. It is appropriately sized for the amount of information offered, though it packs the price tag on with a semicolon in a slightly awkward way.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, but for a tool with three parameters, nested objects, and 0% schema coverage the definition is materially incomplete. It gives no indication of what must go into 'subject' or what qualifies as valid 'evidence', which an agent needs in order to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for three parameters including two nested objects, so the description must compensate but does not. It never mentions 'subject', 'signals', or 'evidence' directly, leaving the agent to infer the shape of the required subject object and the maxItems=32 evidence array from the raw schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific operation (evidence-bounded reputation assessment) on a specific resource (an agent subject), and lists what the result contains (confidence, coverage, provenance limitations, signed receipt). It is clearly distinguishable from generic siblings like verify or attest, though it never names a sibling to differentiate against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance and no alternative tool is named. The phrase 'evidence-bounded' weakly implies evidence must be supplied, but nothing tells the agent which of the many sibling preflight/verify/screening tools to prefer in a given situation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_session_continuityAgent Session ContinuityCRead-onlyIdempotentInspect
Agent Session Continuity; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnlyHint, idempotentHint, and destructiveHint. The description adds 'deterministic' and 'signed,' which are genuine behavioral details beyond the annotations, but it does not describe failure modes, input constraints, or what the signed assessment is used for. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is admirably short and front-loads the core idea, but it is terse to the point of underspecification. The opening phrase repeats the tool title and the remaining clause is generic, so while there is no filler, the brevity sacrifices necessary context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema and annotations cover the return and safety profile, but the description does not explain what session continuity means, what inputs should contain, or when this assessment is needed. Given the required previous/current inputs and a large sibling set, the description is not complete enough for reliable tool selection or invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema lists required fields 'previous' and 'current' without property descriptions, and the description only refers to 'supplied input,' adding no meaning to these fields. Although context reports high schema coverage and zero formal parameters, an agent still has to infer what 'previous' and 'current' represent, and the description does not help.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states that the tool produces a 'deterministic signed assessment over supplied input,' which gives a general action and resource, but it never explains what 'Agent Session Continuity' actually determines or how it relates to the input. The phrase is vague and does not differentiate this tool from sibling tools like signed_result_comparator or verify.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool, what preconditions exist, or when a sibling tool would be more appropriate. The description only defines the tool in isolation and provides no context for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_task_handoff_receiptCInspect
Sign a bounded agent-task handoff; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full behavioral burden. It discloses the cost (0.003 USDC) and that the handoff is 'bounded', which implies a scoping constraint, but says nothing about permission requirements, irreversibility, idempotency, or what happens if payment fails.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no wasted words, and the action verb plus the payment cost are front-loaded. Brevity is efficient here rather than merely under-specified, though it verges on too terse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, but this is a payment-bearing signing operation with zero annotations and the description leaves the key operational context (payment flow, bounded scope, failure handling) unstated. For a tool with monetary and signing implications, that is a significant gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is reported at 100%, so the structured schema is expected to document the required 'handoff' parameter. The description adds no syntax, format, or content guidance for that parameter beyond restating that a handoff is being signed, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Sign') and resource ('bounded agent-task handoff'), which is enough to distinguish it from the many guard/evidence siblings. It does not, however, name any sibling or clarify how it differs from receipt-producing tools like delivery_evidence or retention_deletion_receipt.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No indication of when to call this versus alternatives, no prerequisites, no exclusions. The only usage-adjacent detail is the price, which hints the call is paid but does not say under what circumstances a handoff should be signed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ap2_mandate_bridgeAP2 Mandate BridgeARead-onlyIdempotentInspect
Map mandate fields for review without advertising unsupported AP2 production authorization.; 0.004 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint. The description adds valuable behavioral context beyond these: a cost of 0.004 USDC, determinism, evidence-bounded behavior, and signed output. It also discloses a limitation (no unsupported AP2 production authorization). No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short but structurally awkward. The semicolon after 'authorization.' creates a fragment, and '0.004 USDC. Deterministic, evidence-bounded and signed.' are separate staccato phrases. While concise, the flow hinders readability and could confuse an agent parsing the sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is low-complexity with zero parameters, rich annotations, and an output schema present, so the description need not cover return values or parameter details. It provides cost, determinism, evidence-boundedness, signing, and a scope limitation, which together are reasonably complete for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero named parameters, so the baseline is 4. The schema allows additional properties with minProperties 1, but the description does not need to explain parameter semantics for a tool with no defined params.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('map') and resource ('mandate fields') with a clear purpose ('for review'). The qualifier 'without advertising unsupported AP2 production authorization' clarifies scope, but it does not explicitly differentiate from siblings such as mandate_reserve or mcp_a2a_task_bridge, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives. There are no mentions of sibling tools, conditions, or exclusions. The phrase 'for review' weakly implies a review context, but that is not sufficient to guide an agent's tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
attestCInspect
x402 paid signed JSON attestation; 0.001 USDC on Base.
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | ||
| sources | No | ||
| ttl_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose two genuine behavioral facts: this is a paid operation and the price/chain are 0.001 USDC on Base. It omits whether it is a write/side-effecting operation, whether the attestation is stored or returned, and what signing or auth is needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single front-loaded fragment with no wasted words, which is efficient. However, the brevity crosses into under-specification rather than conciseness, so the structure succeeds only as a one-line label.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a paid, signed operation with three undocumented parameters, no annotations, and no output schema, the description is materially incomplete. An agent knows the price but not what it must pass, what it gets back, or how the result is consumed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all three parameters (content, sources, ttl_seconds), and the description explains none of them. The required 'content' parameter and the ttl_seconds semantics (expiry of the attestation?) are left entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (signed JSON attestation) and the mechanism (x402 paid), which tells an agent what it produces, but it reads as a jargon fragment rather than a clear verb+resource statement. It does not distinguish this from siblings like verify or fetch_evidence, which plausibly operate on the same attestation concept.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus verify, fetch_evidence, or preflight/payment_preflight. The cost mention implies a paid path exists, but the description never says to call payment_preflight first or what condition selects attestation over verification.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
automated_dispute_bundleAutomated Dispute BundleARead-onlyIdempotentInspect
Assemble failure evidence and recommend resolution without executing refunds.; 0.005 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal read-only, idempotent, and non-destructive behavior; the description adds useful non-obvious traits: a deterministic output, an evidence-bounded computation, a cryptographic signature, and a 0.005 USDC cost. These go beyond the structured hints, though 'signed' is not elaborated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with purpose, and every phrase adds signal—scope, non-goal, cost, determinism, boundedness, and signing. The semicolon-separated fragments are slightly awkward but not bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with an output schema and safety annotations, the description covers purpose, non-goal, cost, determinism, and signing. The main remaining gap is the unconstrained input object: minProperties 1 with additionalProperties true leaves agents guessing what at least one property should be.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no declared parameters and the schema only permits a free-form object with at least one property, so the baseline for this dimension is 4. The description's 'failure evidence' phrase hints at the intended payload, though it does not spell out accepted keys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Assemble'), a concrete object ('failure evidence'), and an explicit non-goal ('without executing refunds'), so the tool's role is immediately distinguishable from sibling payment/execution tools. It also names the expected output ('recommend resolution').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case—bundling evidence and recommending a resolution after a failure—and explicitly excludes refund execution, but it names no sibling tools or conditions that should route an agent here instead of related tools like payment_diagnose or transaction_recovery. It leaves the when-to-use decision largely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
autonomous_procurementAutonomous ProcurementARead-onlyIdempotentInspect
Recommend a verified eligible provider under a mandate without purchasing.; 0.005 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly and idempotent annotations, the description discloses a 0.005 USDC cost, determinism, evidence-boundedness, and signed output. These are meaningful behavioral traits that are not present in the annotations or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The content is dense and front-loaded: the core action, scope, cost, and key behavioral properties all appear in a short span. The punctuation is awkward ('without purchasing.; 0.005'), but every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core behavior, cost, and output properties, while the output schema handles return-value details. It could be more explicit about what inputs are expected and what counts as an eligible provider under a mandate, but it is adequate for a zero-parameter read-only tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Context reports zero formal parameters and 100% schema description coverage, so the baseline of 4 applies. The description does not document the open-schema input keys, but with no defined parameters this is not a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Recommend') and a specific resource ('verified eligible provider under a mandate'), and explicitly states the tool does not purchase. This clearly separates it from payment, bridge, and mandate-execution siblings even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context: when a mandate exists and a provider recommendation is needed without purchasing. However, it does not explicitly name alternatives or state when not to use this tool, leaving selection among siblings to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capability_benchmarkCapability BenchmarkBRead-onlyIdempotentInspect
Measure reproducible success, latency and cost from versioned evidence-bound benchmark cases.; 0.004 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive behavior, and the description adds meaningful context: deterministic execution, evidence-boundedness, signed results, and a specific 0.004 USDC cost. This goes beyond the annotations, though the exact meaning of 'evidence-bound' and the pricing fragment remain somewhat underspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loads the core purpose. However, the semicolon-separated '0.004 USDC' fragment and the staccato tag list read as disconnected metadata rather than integrated prose, which makes the structure choppy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With zero parameters and an output schema, the essential invocation details are mostly covered: what is measured and the key guarantees. Still, the description does not explain what actually drives the benchmark, what 'signed' outputs mean, or how it relates to the many neighboring capability and verification tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool reports zero parameters and 100% schema description coverage, so there is no parameter documentation burden. The description hints at implicit benchmark cases, but since no explicit parameters exist, this is not a meaningful gap. The baseline for parameterless tools is appropriately credited.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Measure') and names concrete resources ('reproducible success, latency and cost' from 'versioned evidence-bound benchmark cases'). It is not a tautology and gives a clear sense of what the benchmark computes. However, it does not explicitly distinguish itself from siblings like capability_verification or signed_result_comparator.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given for when to use this tool instead of alternatives. The phrase 'from versioned evidence-bound benchmark cases' implies a prerequisite, but it is not stated as a decision rule. There are no exclusions or references to sibling tools, so an agent receives little routing help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capability_negotiation_preflightCapability Negotiation PreflightCRead-onlyIdempotentInspect
Capability Negotiation Preflight; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds 'deterministic signed assessment', which is useful behavioral context beyond the annotations, but it does not explain what signing entails, what the assessment covers, or what happens to the input. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence with no wasted words. It front-loads the tool name and then adds the key qualifiers 'deterministic signed assessment'. It is concise, though it could be more informative without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (two required params, output schema present, many siblings), the description is too thin. It does not explain what the assessment evaluates, how 'required' and 'offered' relate to capability negotiation, or when an agent should choose this over capability_verification or preflight. The output schema exists but the description still leaves the tool's purpose ambiguous.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the two required parameters ('required' and 'offered'). The description adds no meaning beyond the schema; it does not explain what 'required' and 'offered' mean in the context of capability negotiation. Baseline 3 is appropriate because the schema carries the parameter documentation burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Capability Negotiation Preflight; deterministic signed assessment over supplied input' states a specific verb-ish resource ('assessment') and mentions determinism and signing, but it does not explain what a 'capability negotiation preflight' actually does or what the assessment is about. Among siblings like capability_verification, capability_benchmark, and preflight, the description does not distinguish this tool from those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus siblings such as capability_verification, capability_benchmark, or preflight. The phrase 'over supplied input' implies it takes input and produces an assessment, but there is no explicit context, exclusions, or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capability_verificationCapability VerificationBRead-onlyIdempotentInspect
Separate claimed capabilities from independently evidenced successful demonstrations.; 0.003 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly and idempotent annotations, the description adds useful traits: deterministic, evidence-bounded, and signed, plus an explicit cost of 0.003 USDC. These are meaningful behavioral signals, though 'evidence-bounded' could be more precise.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, with the purpose and key constraints front-loaded. Minor punctuation quirks aside, every clause adds information without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema, the description omits crucial context: what the input object expects (the schema allows any additional properties), what 'evidence-bounded' means in practice, and whether the 0.003 USDC cost requires a transaction beforehand. The description is too abstract for agents to safely invoke.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool declares zero parameters, and the schema coverage is 100%, so the baseline of 4 applies. No parameter documentation is needed, and the description adds nothing that could be misread as parameter guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear goal—'separate claimed capabilities from independently evidenced successful demonstrations'—but it reads more like an outcome than a concrete tool action with a specific verb and resource. The title and sibling names hint at verification, yet the exact operational function remains ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given for when to use this tool versus alternatives like capability_benchmark, capability_negotiation_preflight, or tool_capability_drift_monitor. There are no use conditions, exclusions, or comparisons, leaving the routing decision to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
company_enrichment_evidenceCInspect
Company enrichment with free input preflight, provider attribution and signed evidence; 0.005 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| lid | No | ||
| name | No | ||
| No | |||
| input | No | ||
| phone | No | ||
| query | No | ||
| entity | No | ||
| pdl_id | No | ||
| person | No | ||
| region | No | ||
| school | No | ||
| ticker | No | ||
| company | No | Company name or nested input object | |
| contact | No | ||
| country | No | ||
| profile | No | ||
| subject | No | ||
| website | No | ||
| locality | No | ||
| location | No | ||
| last_name | No | ||
| birth_date | No | ||
| email_hash | No | ||
| first_name | No | ||
| identifier | No | Email, profile URL, phone, company domain, ticker or name; interpreted by service | |
| postal_code | No | ||
| street_address | No | ||
| minimum_likelihood | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does add meaningful context by disclosing a free input preflight, provider attribution, signed evidence, and a fixed price of 0.005 USDC, which helps an agent understand billing and validation behavior. However, it omits broader operational details such as authentication requirements, data source behavior, failure modes, and return structure, so it is only partially transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence with no filler: it names the operation, key features, and cost. The information is packed tightly, though it slightly sacrifices clarity by relying on jargon like 'provider attribution' without elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 28 parameters, nested objects, no output schema, and no annotations, the description is too thin to support correct invocation. It provides no invocation guidance, no examples, no explanation of what the signed evidence contains, and no comparison to similar enrichment tools. An agent needing to choose parameters or interpret the result would be under-informed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 7%, and the tool description itself adds no parameter-level meaning beyond 'company enrichment.' The schema's own description does clarify that identifier can be pdl_id, website/domain, profile, ticker, name, company, or identifier and that nested inputs are accepted, but the remaining ~26 parameters are opaque. The description does not compensate for the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a clear domain ('Company enrichment') and distinguishes it from sibling evidence tools such as contact_enrichment_evidence, person_enrichment_evidence, and entity_enrichment_evidence by naming the company scope. It does not state what the returned 'signed evidence' contains, which leaves some ambiguity about the tool's exact output, but the primary purpose is recognizable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use this tool versus sibling alternatives like contact_enrichment_evidence or entity_enrichment_evidence. There are no exclusion criteria, no mention of alternatives, and no condition-based routing. The only implicit signal is that it is for company-related enrichment.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
concurrency_guardConcurrency GuardCRead-onlyIdempotentInspect
Fail closed unless capacity counter is fresh and the request carries a lease plus idempotency binding; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| policy | Yes | ||
| lease_id | Yes | ||
| active_slots | Yes | ||
| idempotency_key | Yes | ||
| requested_slots | Yes | ||
| lease_valid_until | Yes | ||
| counter_observed_at | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description usefully adds the fail-closed posture, the freshness requirement on the counter, and the 0.002 USDC cost, which annotations do not convey. It stops short of saying what a denial looks like or whether the call itself performs the slot reservation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the failure condition and appends the price. Nothing is wasted, though the compression is extreme enough to obscure the tool's core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-required-parameter guard with no output schema and no field descriptions, the definition is thin: it does not explain how active_slots plus requested_slots are compared against policy.maximum_slots, nor what the caller receives on pass versus fail-closed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 7 required parameters, so the description carries the full burden. It loosely hints at three parameter clusters (counter freshness → counter_observed_at, lease → lease_id/lease_valid_until, idempotency binding → idempotency_key) but leaves requested_slots, active_slots, and policy.maximum_slots unexplained, and never states the freshness window or timestamp units.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a decision policy ('fail closed unless...') and the inputs that gate it (fresh counter, lease, idempotency binding) plus a price, so the resource is identifiable. However, it never states the tool's overall function in plain terms (e.g., admit/deny a request against a concurrency cap), and it does not distinguish itself from close siblings such as idempotency_replay_guard or resource_cycle_guard.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no named alternative. Given the dense set of guard siblings (idempotency_replay_guard, resource_cycle_guard, delegation_scope_guard), the agent has no signal for choosing this one over them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
conflict_resolutionCInspect
Resolve or abstain; every candidate must bind source, value hash, observation time and bounded confidence; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| policy | Yes | ||
| candidates | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It does add genuinely useful behavior: an abstain path, a bounded-confidence constraint, and a concrete price of 0.003 USDC. It omits permissions/payment mechanics, what happens when candidates are irreconcilable, and the result shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the operation and appends the per-candidate requirements and price. No filler, though the semicolon-separated clauses make it slightly terse to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with a nested array, a nested policy object, no output schema and no annotations, the description is too thin. It never explains the policy structure, the resolution rules, or the return value, leaving an agent unable to invoke it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and objects are nested, so the schema explains little. The description usefully enumerates the required candidate bindings (source, value hash, observation time, bounded confidence), which covers the candidates parameter, but the required 'policy' object is left entirely opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The verb pair 'Resolve or abstain' combined with the name gives a rough idea that this adjudicates conflicting candidates, and it names the shape candidates must take. But it never states the resource being resolved (conflicting values? sources?) or what resolution produces, so the purpose stays implied rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to call this versus siblings like signed_result_comparator or verify, nor any precondition or exclusion. 'Resolve or abstain' hints at output behavior, not usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contact_enrichment_evidenceCInspect
Business-contact enrichment with free input preflight, identity limitations and signed receipt; 0.007 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| lid | No | ||
| name | No | ||
| No | |||
| input | No | ||
| phone | No | ||
| query | No | ||
| entity | No | ||
| pdl_id | No | ||
| person | No | ||
| region | No | ||
| school | No | ||
| ticker | No | ||
| company | No | Company name or nested input object | |
| contact | No | ||
| country | No | ||
| profile | No | ||
| subject | No | ||
| website | No | ||
| locality | No | ||
| location | No | ||
| last_name | No | ||
| birth_date | No | ||
| email_hash | No | ||
| first_name | No | ||
| identifier | No | Email, profile URL, phone, company domain, ticker or name; interpreted by service | |
| postal_code | No | ||
| street_address | No | ||
| minimum_likelihood | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does disclose useful behavioral traits: free input preflight, identity limitations, signed receipt, and the 0.007 USDC cost. But it omits whether the operation is read-only, what the signed receipt contains, and possible failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single tight sentence, front-loading the core purpose and packing cost, preflight, and identity constraints without filler. It could be slightly better structured, but it earns its brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 28 parameters, no output schema, and no annotations, one line is insufficient. An agent cannot determine which identifier to pass, what the signed receipt looks like, or how identity limitations affect the enrichment result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 7%, yet the description names no parameters and gives no guidance on the 28 inputs. It only implies the input is a business contact, which does little to clarify how to populate fields like lid, pdl_id, query, or subject.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Business-contact enrichment' – a specific verb and resource – and the 'business-contact' qualifier helps distinguish it from company or person enrichment siblings. However, it does not name any sibling or explicitly state what the evidence output looks like.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No usage context or alternatives are provided. The description mentions preflight and identity limitations but never says when to use this tool over company_enrichment_evidence, person_enrichment_evidence, or entity_enrichment_evidence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
context_provenance_labelerCInspect
Hash and label context integrity, confidentiality and source; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses the 0.002 USDC charge, but says nothing about what 'hash and label' produces, whether the hashed content is transmitted or retained, or whether the operation is deterministic — significant gaps for a tool that ingests arbitrary content.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single compact sentence with the core action front-loaded and the cost appended. Nothing is wasted, though the terse phrasing leans on jargon that would benefit from a short clarifier.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. However, with no annotations and no usage context for a content-ingesting, paid operation, the definition leaves the agent guessing about safety, retention, and when this tool is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is reported at 100%, so the schema is presumed to document 'content' and 'source' adequately, giving a baseline of 3. The description adds no meaning beyond that, such as expected content types or source format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names specific verbs (hash, label) and the resource (context integrity, confidentiality and source), so an agent can tell it produces a provenance/integrity artifact. It does not, however, distinguish itself from siblings in the same evidence/attestation family (e.g. dependency_provenance_assessment, attest, fetch_evidence).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives among the numerous sibling guard/evidence tools. The only decision signal is the price, which tells the agent it costs money but not when it is the right tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
counterparty_risk_preflightCInspect
Deterministic counterparty policy preflight using supplied reputation evidence and transaction limits; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| policy | No | ||
| subject | Yes | ||
| reputation | No | ||
| transaction | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose two useful traits beyond the schema: the evaluation is deterministic (same inputs yield the same verdict) and it costs 0.003 USDC. It omits whether the call is read-only, whether it requires auth or holds/reserves funds, and what happens on a failing verdict.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the core capability first and the price appended. Nothing is padded, though the trailing cost fragment is slightly awkward rather than fully integrated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but the tool has four undocumented nested-object inputs and no annotations. For a deterministic policy-evaluation tool, the description leaves the agent unable to construct the required 'subject' and 'policy' payloads.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across four nested-object parameters, so the description must compensate and largely does not. It only hints at two of them ('reputation evidence', 'transaction limits') and says nothing about the 'policy' or 'subject' objects or the shape of any nested field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (counterparty policy) and operation (deterministic preflight) and states what inputs drive it (reputation evidence, transaction limits). It is clear enough to distinguish from generic siblings like 'verify' or 'try_service', but it never explicitly contrasts itself with the closely related 'payment_preflight' or 'preflight' tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this versus the many preflight/assurance siblings (payment_preflight, preflight, transaction_assurance). No prerequisites, ordering, or exclusions are given; the agent must infer that 'preflight' means 'call before transacting' from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cross_agent_receipt_bundleCross-Agent Receipt BundleCRead-onlyIdempotentInspect
Cross-Agent Receipt Bundle; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive behavior; the description adds that the result is deterministic and signed, which is a genuine extra behavioral trait. However, it does not explain the signature's role, failure modes, or assessment semantics, so the contribution is modest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short, but the opening clause is a verbatim repeat of the title and the remaining clause is too generic to carry meaningful information. This reads as under-specification rather than disciplined conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a required `receipts` input, a large sibling tool set, and no prose guidance, an agent cannot reliably determine what to pass or when to select this tool. The output schema and annotations cover some safety/return expectations, but the core invocation and routing questions are entirely unanswered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema names a required `receipts` field but provides no property description, and the phrase 'supplied input' is a placeholder that adds no meaning. The 100% schema-description coverage does not help because the schema itself contains no real parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first clause simply restates the tool's title. 'Deterministic signed assessment over supplied input' names a generic action but not what the tool actually bundles, assesses, or returns, and it does not distinguish this from the many sibling receipt/bundle/evidence tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given for when to use this tool versus alternatives such as agent_task_handoff_receipt, cross_protocol_receipt, or multi_source_fact_bundle. There are no conditions, prerequisites, or exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cross_protocol_receiptCross-Protocol ReceiptBRead-onlyIdempotentInspect
Bundle consistent transaction commitments across multiple protocols.; 0.004 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint. The description adds meaningful behavioral context beyond those hints: 'Deterministic, evidence-bounded and signed' plus the explicit cost of '0.004 USDC'. This enriches the agent's understanding of side effects and execution guarantees without contradicting annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and front-loaded with the core function. The cost and property descriptors add value without padding. The semicolon/dangling '0.004 USDC' fragment is slightly awkward structurally, but overall the text is efficiently writed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite good annotations and an output schema, the description is too thin for a tool with an open input schema (minProperties 1, additionalProperties true). It does not explain what transaction commitments should be supplied, what 'consistent' means in practice, how protocols are identified, or how the bundling is requested. An agent has little concrete guidance for constructing a valid invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema declares no named parameters (0 params, 100% schema description coverage), so the baseline is 4. The description does not need to explain parameter meaning because there are no explicit parameters to document; however, it also does not clarify what should go into the allowed arbitrary object, which is a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Bundle') and resource ('consistent transaction commitments across multiple protocols'), clearly conveying the core function. It is not a tautology and the cross-protocol qualifier distinguishes it from generic receipt tools, though it does not explicitly name a sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as cross_agent_receipt_bundle or automated_dispute_bundle. No prerequisites, exclusions, or selection criteria are provided; usage must be inferred entirely from the purpose statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
data_egress_preflightBInspect
Inspect outbound agent data and destination before transmission; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| policy | No | ||
| payload | Yes | ||
| destination | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full behavioral burden. It does disclose a genuinely useful trait beyond the schema: a per-call cost of 0.002 USDC. But it omits whether inspection is advisory or blocking, what the policy argument changes, and any auth or rate-limit behavior for a paid gate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded clause with no filler; the scope and cost are stated compactly. It is efficient, though the extreme brevity is part of why the behavioral and parameter gaps exist.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, but for a paid, complex-input preflight with zero annotation coverage and 0% parameter documentation, the description leaves key questions unanswered (blocking vs advisory, policy semantics, payload format). It is too thin for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across three parameters, including two required nested/opaque objects. The description conceptually names 'data' and 'destination' (mapping loosely to payload and destination) but says nothing about the optional 'policy' object, what shape payload/destination take, or what policy controls.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Inspect') with its resources ('outbound agent data and destination') and timing scope ('before transmission'), which is enough to distinguish it from payment-oriented siblings like payment_preflight. It stops short of explicitly naming which sibling to use instead, so it falls just below full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Before transmission' implies the usage context, so an agent can infer this is a pre-send gate. However, with numerous preflight/guard siblings (counterparty_risk_preflight, tool_call_policy_guard, transaction_assurance), no guidance is given on when this is the correct gate versus the others, and no exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
data_freshness_certificateBInspect
Certify freshness only when timestamp is bound to source URI and a 64-hex content hash; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| data | Yes | ||
| policy | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses a hard prerequisite (the binding requirement) and the cost (0.002 USDC), and the phrase 'only when' implies refusal otherwise. But it says nothing about failure behavior, what the certificate contains, or how the maximum_age_seconds policy is enforced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the core verb and constraints, with cost at the end. No wasted words, though the compression leaves semantic gaps that a slightly longer definition could close.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with nested required schemas, zero parameter descriptions, no annotations, and no output schema, the single sentence is not sufficient. It omits the policy semantics, the certificate's contents, and the failure/return behavior an agent needs to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and both parameters are nested required objects. The description maps three data subfields (observed_at, source_uri, content_sha256) to their intent and adds the '64-hex' format detail, which is genuinely beyond the schema. But the entire policy object, including the central maximum_age_seconds threshold, is left unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (Certify) and resource (freshness) with the exact binding conditions required to issue the certification. It is clearly distinguishable from generic siblings like attest or verify by the freshness/URI/hash semantics, though it does not explicitly name a sibling it is not. Purpose is clear but sibling differentiation is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a precondition (only when timestamp is bound to source URI and a 64-hex content hash), which is a form of when-to-use guidance. However, it does not explain when to prefer this over nearby tools like attest, verify, or service_sla_attestation, nor any exclusions. Usage is implied rather than routed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delegated_credential_guardDelegated Credential GuardCRead-onlyIdempotentInspect
Delegated Credential Guard; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose read-only, idempotent, and non-destructive behavior. The description adds that the result is deterministic and signed, which is useful but minimal; it does not explain what input is accepted, what the assessment contains, or whether any external state is involved.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and wastes few words, and the key functional hint 'deterministic signed assessment' is front-loaded. However, the leading phrase 'Delegated Credential Guard' simply duplicates the tool name and adds no information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema and annotations, the description leaves the tool's core semantics undefined: what a delegated credential assessment is, what inputs are expected, what output shape or meaning results, and when an agent should prefer it over similar guards. This is too sparse for safe autonomous invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema names two required inputs, `credential` and `action`, but provides no property definitions or descriptions. The description does nothing to clarify what format these should take, what 'action' means, or how they relate to the assessment, so an agent cannot reliably construct a valid call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description merely pairs the tool name with a generic phrase, 'deterministic signed assessment over supplied input,' without stating what kind of credential or action is assessed, what the guard actually does, or what makes it different from the many sibling guard/assessment tools. It restates the title and adds only vague functional language.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool, when not to use it, or which sibling tools are alternatives. The description gives an agent no contextual signal to distinguish delegated_credential_guard from delegated_spend_policy, delegation_scope_guard, or attest.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delegated_spend_policyDelegated Spend PolicyBRead-onlyIdempotentInspect
Fail closed when delegated authority, per-call limit or remaining budget is insufficient.; 0.003 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly, idempotent, and non-destructive behavior, and the description adds meaningful context: fail-closed semantics, deterministic results, evidence-bounded output, and signed results, plus an apparent 0.003 USDC cost. This goes beyond the structured fields without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The text is short and front-loaded with the core failure behavior, but the remaining clauses are sentence fragments; '0.003 USDC' is ambiguous between a cost and a threshold, and the list grammar is slightly off. It earns points for brevity but loses polish.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter policy tool with an output schema and safety annotations, the description covers the core decision, determinism, and signed evidence. It is incomplete on when to choose this tool over related guards and on what exact input object the tool expects, which matters because the schema is intentionally open.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero documented parameters and the input schema is an open object, so the description must carry the semantic weight; it names the relevant conceptual inputs (delegated authority, per-call limit, remaining budget). It still doesn't provide exact keys or value formats, but with no schema-defined parameters this is a reasonable baseline-4 case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first clause names the resource (delegated spend) and the core behavior (fail closed when delegated authority, per-call limit, or remaining budget is insufficient), so an agent can tell this is an authorization/policy gate. It does not explicitly state a success-path action or differentiate itself from the many guard-style sibling tools, which keeps it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no 'use when' or 'instead of' guidance, and none of the many sibling guard tools are named as alternatives. The only usage cue is the failure condition, which tells an agent when the tool would block, not when it should be called.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delegation_scope_guardCInspect
Least-privilege guard for agent-to-agent delegation scope, destinations, budget and expiry; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| policy | Yes | ||
| delegation | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears the full burden of behavioral disclosure. It does not state whether the guard blocks, warns, or logs, what happens when a policy is violated, what the output contains, or what the 2 required nested objects must include. Only the protected dimensions are listed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence with no filler. It front-loads the guard's purpose, though it omits any structural cue about the required objects.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No annotations, no output schema, two required nested objects with 0% schema description coverage, and no behavioral details. The description is far too thin for a tool that appears to make security-critical decisions about delegation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and both parameters are required nested objects with no shape or property descriptions. The description adds no meaning beyond naming the domains, so an agent cannot infer what 'delegation' and 'policy' must contain.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action: enforcing 'least-privilege guard' on 'agent-to-agent delegation scope, destinations, budget and expiry'. However, no verb directly states what the tool does (e.g., validate, check, enforce). It also lacks distinguishing context from siblings like payment_preflight or counterparty_risk_preflight, which also appear to be pre-execution guards.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance. The description does not explain in which scenarios an agent should invoke this guard versus alternatives such as preflight or counterparty_risk_preflight.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delivery_evidenceDelivery EvidenceCRead-onlyIdempotentInspect
Bind successful delivery to transaction, request and matching 64-hex expected/observed content hashes without echoing content; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| delivery | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is lower. The description adds genuinely useful behavior: it does not echo content, and it requires matching 64-hex expected/observed hashes. But it omits what counts as 'successful', what a mismatch returns, and the pricing/billing mechanics beyond the flat fee.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence followed by the price. Front-loaded with the core action and no filler. The dense clause stacking slightly hurts readability but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A paid, mutation-adjacent evidence tool with a nested 6-field object, 0% schema coverage, and no output schema. The description should explain the input object shape, the matching rule, mismatch behavior, and the return value. It covers only the hash-matching idea and the price, leaving most of what an agent needs to call it correctly unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single required 'delivery' parameter is a nested object with six required sub-fields, none documented. The description gestures at the fields (transaction, request, content hashes) but supplies no format, no max length enforcement, and no explanation of the hash-matching semantics the caller must satisfy.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a verb ('Bind') and the resources involved (delivery, transaction, request, content hashes), which is more specific than a tautology. However 'Bind successful delivery' is jargon that doesn't clearly explain to an agent what the tool actually produces or how it differs from near siblings like transaction_assurance, agent_task_handoff_receipt, or attest. The purpose is inferred rather than stated plainly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no mention of alternatives, despite dozens of closely-named evidence/attestation siblings. The '0.003 USDC' price hints at a paid operation but the description never says when this paid call is warranted versus free siblings like schema_normalize_free or try_service.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dependency_provenance_assessmentCInspect
Fail closed on empty dependency sets or absent registry policy; distinguish digest declarations from verified provenance; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| policy | Yes | ||
| dependencies | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose meaningful behavior: it fails closed on empty dependency sets or missing registry policy, and it separates claimed digests from verified provenance. It also discloses metered pricing (0.003 USDC). Still missing are whether it performs network/registry lookups, what an assessment result looks like, or any auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact clauses, front-loaded with the fail-closed constraint and ending with the cost, with no padding or repetition. It is efficient, though the brevity contributes to under-specification elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two required parameters, a nested object schema at 0% description coverage, no output schema and no annotations, the description is far too thin to call correctly. It neither explains the input contract nor what the assessment yields, leaving core questions unanswered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and both parameters (including a nested policy object with allowed_registries) are undocumented at any level. The description only implicitly alludes to 'dependency sets' and 'registry policy' without specifying entry format, registry identifier shape, or the 1000-item cap, so it does almost nothing to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The name and phrase 'distinguish digest declarations from verified provenance' gesture at the function, implying a supply-chain provenance check against registry policy. However, the description never states plainly what the tool does or returns (e.g., 'assesses/attests dependency provenance'), and it offers no differentiation from nearby siblings like fetch_evidence or mcp_manifest_firewall. Purpose is inferable but vague.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool, what prerequisites exist, or which siblings are alternatives. The clause 'fail closed on empty dependency sets or absent registry policy' describes a guard behavior, not a usage condition that tells an agent when to select this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
document_to_verified_jsonCInspect
Public text, HTML, JSON or XML normalized to source-backed JSON; no OCR; 0.010 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| source_urls | Yes | ||
| expected_fields | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It usefully discloses accepted input formats, the source-backed nature of the output, the lack of OCR support, and the price. However, it does not describe failure modes, rate limits, processing guarantees, or whether any action is taken beyond normalization.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and front-loaded, with the core transformation stated first, followed by the OCR limitation and cost. Every clause adds useful information. It is slightly too terse given the unexplained expected_fields parameter, but the structure itself is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has two parameters, no annotations, no output schema, and only a one-sentence description. The description fails to explain expected_fields, input URL semantics, output shape, or behavioral constraints beyond OCR. An agent cannot confidently invoke this tool correctly from the provided information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must clarify the parameters. It indirectly describes source_urls as public text/HTML/JSON/XML sources, but it does not explain expected_fields at all, and gives no hint about their format, purpose, or relationship to the output. This leaves a required part of the contract unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the resource (public text, HTML, JSON, XML) and the result (source-backed JSON). It positions the tool as a normalizer/converter and adds a differentiating constraint ('no OCR'). It does not explicitly distinguish it from siblings such as source_backed_search or verified_web_extract, but the core purpose is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Public text, HTML, JSON or XML' and 'no OCR' imply the tool is for public non-image document conversion, but there is no explicit statement of when to use this tool versus a sibling, and no mention of when not to use it. No alternatives or routing conditions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
domain_ownership_evidenceCInspect
Require a fresh domain-bound DNS-01/HTTP-01 challenge and separate control from legal ownership; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | ||
| evidence | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose the price (0.003 USDC), the freshness requirement, the accepted proof methods, and the control-vs-legal-ownership distinction — useful behavioral context. It omits whether the call is read-only or mutating, what happens on failure, and what the caller receives.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with no filler, and the cost and freshness constraint are placed up front. The phrasing 'Require a fresh ... challenge' is elliptical but compact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations, no output schema, and an undocumented nested evidence object, the description leaves too much for the agent to infer about inputs and outcomes. Given the tool's structural complexity, substantially more detail was warranted.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two parameters, one of which is a nested object with five required subfields. The description maps loosely onto 'domain' and, via DNS-01/HTTP-01, onto the evidence method, but leaves challenge_sha256, verified_at, and independently_verified unexplained, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a domain-ownership verification via DNS-01/HTTP-01 challenges and distinguishes control from legal ownership, which is a meaningful resource-specific idea. However, the leading verb 'Require' is an odd framing for a tool call and never states plainly what the tool does (verify, attest, issue evidence, or gate an action). An agent knows roughly the subject matter but not the operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of alternatives among the many evidence sibling tools (rwa_sanctions_screening_evidence, agent_reputation_evidence, mcp_server_identity_evidence, etc.), and no preconditions stated beyond a fresh challenge. The 'freshness' hint is the only contextual cue.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
duplicate_charge_detectorDuplicate Charge DetectorBRead-onlyIdempotentInspect
Detect repeated transaction hashes, payment identifiers and idempotent intents; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| charges | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the agent knows this is a safe, pure, repeatable read operation. The description adds the cost (0.003 USDC), which is useful operational context not present in annotations. However, it does not describe output format, detection sensitivity, or how duplicates are identified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the purpose and appends the cost. It wastes no words, though it could be slightly more structured by separating purpose and cost.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only detection tool with no output schema, the description covers what is detected and the cost. However, it lacks details on input format, output shape, and when to use it, which are needed for correct invocation given the 0% schema description coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the burden of explaining the 'charges' parameter. The description mentions 'transaction hashes, payment identifiers and idempotent intents', which hints at what the charges array should contain, but does not specify the expected object shape, required fields, or limits (maxItems: 1000 is only in the schema). This leaves a significant gap for the sole required parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (detect) and resource (duplicate/repeated transaction hashes, payment identifiers, idempotent intents). It clearly distinguishes itself from adjacent siblings like payment_preflight and idempotency_replay_guard by focusing specifically on duplicate detection. However, it does not explicitly differentiate itself from payment_diagnose or payment_receipt_reconciler.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives. The description does not mention prerequisites, when-not-to-use cases, or which sibling tools handle related concerns. The only contextual cue is the cost, which helps with selection but is not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
durable_agent_taskDurable Agent TaskCRead-onlyIdempotentInspect
Durable Agent Task; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds behavioral context with 'deterministic signed', implying output is reproducible and cryptographically signed, which goes beyond annotations. However, it does not explain what 'signed' entails, whether authentication is needed, or what the output format is. With annotations doing most of the work, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and front-loaded with the tool name, which is efficient. However, 'Durable Agent Task' simply repeats the title and does not earn its place. The rest of the sentence adds some value, so overall it is mostly concise but with a minor redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is too vague for a tool with one required parameter and no parameter descriptions. It does not explain what a 'durable agent task' is, what 'deterministic signed assessment' means for the caller, or how the 'task' input should be provided. Although an output schema exists and annotations cover safety, the description does not give enough context for an agent to call this correctly, especially given many similar sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is reported as 100%, so the baseline is 3. The description adds 'supplied input', which suggests the 'task' parameter is the input being assessed, but it does not clarify the expected structure, format, or constraints of that input. This provides minimal additional meaning beyond the parameter name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description says it is a 'deterministic signed assessment over supplied input', which gives a rough sense of the tool's function but lacks a specific verb or clear resource. It does not distinguish itself from sibling tools like 'attest' or 'verify', which could also be described as signed assessments. It is more than a tautology but remains vague about what the tool actually does (executes a task? returns an assessment? creates a durable record?).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool or how it differs from alternatives. No conditions, exclusions, or references to sibling tools are provided. An agent would have no way to decide between this and similar-sounding tools like 'attest' or 'signed_result_comparator'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dynamic_service_pricingDynamic Service PricingBRead-onlyIdempotentInspect
Propose a capped shadow price without changing the live catalog.; 0.004 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, and non-destructive behavior. The description adds meaningful context: deterministic, evidence-bounded, signed, and no live-catalog change. However, it leaves 'signed' and 'evidence-bounded' undefined, and the fee '0.004 USDC' is unclear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very brief and front-loaded with the core purpose, followed by three short constraints. The stray '0.004 USDC.' feels abrupt and there is a typo ('catalog.;'), but every sentence carries intent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only tool with an output schema, the description is nearly sufficient. Yet it doesn't clarify whether the caller should supply arbitrary properties (schema allows additionalProperties) or what 'evidence-bounded' means in practical terms, leaving some ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero defined parameters, and the schema is a generic object with additionalProperties true. The description adds no parameter-level guidance, but with zero formal parameters the schema inherently covers everything. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (propose) on a specific artifact (capped shadow price) and explicitly notes it does not modify the live catalog, which distinguishes it from price-update tools. However, the phrase 'capped shadow price' is domain-specific and not fully self-explanatory to an unfamiliar agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a use case (proposing a price without changing the catalog) but provides no explicit when-to-use guidance, prerequisites, or alternatives like price_discovery_engine. An agent would have to infer when this tool is appropriate relative to its many siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
economic_loop_detectorEconomic Loop DetectorBRead-onlyIdempotentInspect
Detect self-transfer and reciprocal economic patterns without presenting correlation as fraud.; 0.004 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint. The description adds meaningful behavioral context: deterministic, evidence-bounded, signed, and a 0.004 USDC cost. There is no contradiction with the read-only annotation since 'Detect' is consistent with it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and front-loads the core purpose, but the semicolon splice before '0.004 USDC' is awkward and the cost fragment is not labeled as a cost. The behavioral terms are high-signal, but the structure feels fragmented rather than polished.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only tool with an output schema, the description covers purpose, cost, determinism, evidence-boundedness, and signing. However, it lacks when-to-use guidance relative to sibling tools and leaves the meaning of 'evidence-bounded' and the '0.004 USDC' fragment unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero named parameters and the schema only accepts an arbitrary object with minProperties 1. With no declared parameter contract, the description cannot meaningfully enrich parameter semantics, and the 0-param baseline of 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific action and resource: 'Detect self-transfer and reciprocal economic patterns'. The clause 'without presenting correlation as fraud' adds useful intent, but it does not explicitly differentiate this tool from siblings like trust_anomaly_detector or sybil_reputation_guard.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus alternatives, and no sibling/alternative is named. The non-fraud caveat is a behavioral guardrail rather than a usage directive, leaving the agent to infer the appropriate context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
entity_enrichment_evidenceCInspect
Source-bounded entity field coverage with explicit missing facts; 0.008 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| subject | No | ||
| source_urls | Yes | ||
| expected_fields | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears the full burden of behavioral transparency. It adds some useful traits—source-bounded results, explicit missing facts, and a 0.008 USDC cost—but it does not disclose read/write behavior, whether it fetches the source URLs, persistence, auth, response shape, or failure semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is compact and front-loaded: scope appears first, missing-facts behavior second, and price last. There is no filler and each clause earns its place. It is under-specified overall, but the structural efficiency is strong.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given three parameters, a nested subject object, no output schema, and a crowded sibling list, the one-sentence description is insufficient. The agent cannot learn how evidence is returned, how subject relates to expected_fields, or when to select this tool over similar ones.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage for three parameters, so the description must compensate. It indirectly maps 'source-bounded' to source_urls, 'entity' to subject, and 'field coverage ... missing facts' to expected_fields, adding some meaning beyond the bare schema names. However, it does not explain formats, requiredness, constraints, or how subject and expected_fields interact.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description conveys a specific output concept—'Source-bounded entity field coverage with explicit missing facts'—and mentions the price, but it uses a noun phrase rather than a clear verb+resource statement. It does not distinguish this from siblings like multi_source_fact_bundle or source_backed_search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no exclusions, and no reference to alternatives. The description only characterizes the result and cost, leaving the agent to infer when this evidence tool should be chosen over the many related tools in the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
erc8004_identity_evidenceERC-8004 Identity EvidenceARead-onlyIdempotentInspect
Bind an ERC-8004 registry identity to evidence while separating registration from ownership.; 0.004 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds valuable behavioral context beyond these: 'Deterministic, evidence-bounded and signed,' plus a 0.004 USDC cost. This enriches the agent's model of the operation without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loads the action. The semicolon leading into '0.004 USDC' is slightly abrupt, and 'separating registration from ownership' is a bit cryptic, but no sentence is wasted and the core behavior is stated first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With output schema present and annotations covering safety and idempotency, the description does not need to explain return values. However, the input schema is open with minProperties=1, and the description does not clarify what properties an agent should supply (e.g., how to specify identity or evidence). For a specialized ERC-8004 tool, this leaves a meaningful gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 declared parameters, so the baseline is 4. The description does not document parameter details, but there are none in the schema to explain. The open additionalProperties=true schema means the description's references to 'identity' and 'evidence' are the only semantic clues, which is acceptable for a zero-param tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Bind an ERC-8004 registry identity to evidence.' This clearly identifies the tool's core function and adds a distinguishing constraint ('separating registration from ownership'). It does not explicitly contrast with sibling tools like erc8004_reputation_intelligence or onchain_evidence, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives. It mentions the cost and behavioral properties but provides no selection criteria, prerequisites, or exclusions, leaving an agent to infer suitability from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
erc8004_reputation_intelligenceERC-8004 Reputation IntelligenceARead-onlyIdempotentInspect
Aggregate diverse verified reputation observations with low-sample abstention.; 0.004 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds meaningful context beyond that: deterministic execution, evidence-bounded results, signed output, low-sample abstention, and a cost of 0.004 USDC. These are behavioral traits an agent could not infer from the annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and front-loads the core purpose. The fragments about cost, determinism, and signed output each add useful information. Minor punctuation oddity ('abstention.;') slightly detracts, but the overall structure is appropriately terse and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Annotations cover the safety profile and an output schema exists, so the description does not need to explain return values. However, the input contract is unclear: the schema allows arbitrary additional properties and minProperties is 1, yet the description never explains what input the tool expects or how it relates to other ERC-8004 tools. This leaves an agent uncertain about what to pass when invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero documented parameters and 100% schema coverage, there are no parameter details for the description to clarify. The baseline for a zero-parameter tool is 4. The open input schema with additionalProperties remains ambiguous, but the description was not required to compensate for undocumented parameters because none are defined.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific operation ('aggregate') and a specific resource ('diverse verified reputation observations'), with a notable behavioral qualifier ('low-sample abstention'). It is not a tautology and gives an agent a working sense of the tool, though it does not explicitly differentiate itself from sibling tools like erc8004_identity_evidence or agent_reputation_evidence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to choose this tool over alternatives, nor does it mention any related sibling tools or exclusions. 'Low-sample abstention' hints at a condition, but there is no practical routing information such as 'use this when you need aggregated evidence' or 'use agent_reputation_evidence for individual observations.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
execution_cost_estimatorExecution Cost EstimatorCRead-onlyIdempotentInspect
Execution Cost Estimator; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the description does not need to restate safety. It adds the behavioral traits 'deterministic' and 'signed' for the assessment, which are useful beyond the annotations, though it does not explain what 'signed' means or mention any other operational behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
At one sentence, the description is compact and the 'deterministic signed assessment' clause is meaningful. However, the opening repeats the title and provides little structured information, making this conciseness more under-specification than efficient clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The schema lists required fields (steps, budget_atomic) without property definitions, so the description needed to explain what input is expected and what kind of assessment is produced. It only says 'supplied input,' leaving important context to inference despite the presence of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema exposes no property definitions (parameter count 0), so the description carries no parameter-semantics burden under the 0-param baseline. The description does not add detail about the required steps and budget_atomic fields, but it also does not mislead; coverage is reported at 100%.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description leads with a restatement of the title ('Execution Cost Estimator') and then describes the tool only as a 'deterministic signed assessment over supplied input.' This lacks a specific verb and resource, and it does nothing to distinguish the tool from siblings such as transaction_risk_score or signed_result_comparator.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no indication of when to use this tool, what it should not be used for, or which sibling alternatives exist. There are no prerequisites, no exclusion conditions, and no reference to the many cost/risk/budget siblings in the list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetch_evidenceBRead-onlyIdempotentInspect
Independently fetch bounded public evidence with redirect/DNS/connected-address SSRF checks, hashes and injection screening; 0.004 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavioral context: it performs SSRF checks on redirects, DNS, and connected addresses, plus hashing and injection screening. Annotations already declare readOnly, openWorld, idempotent, and non-destructive. The cost mention is helpful, but there's no detail on failure modes, rate limits, or output format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single efficient sentence that packs in key features and cost. Front-loaded with the core action. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and a single parameter, the description covers the core action and security checks. However, it omits return value details, error handling, and cost implications beyond the flat fee. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there is one required parameter 'url'. The description doesn't explain the parameter's expected format beyond the schema's 'uri' type, but 'bounded' and 'public' imply restrictions on what URLs are allowed. Baseline 3 when schema has minimal elaboration.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'fetch bounded public evidence' with the added detail of SSRF checks and hashing. It's clear what the tool does, though it doesn't explicitly contrast with siblings like 'try_service' or 'webhook_verifier'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No indication of when to use this tool versus alternatives. The '0.004 USDC' cost hints at a paid service, but there's no guidance on when it's appropriate or what happens if the fetch fails.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
human_approval_policyBInspect
Classify an action as automatic, approval-required or denied; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears the full behavioral burden. It usefully discloses a cost of 0.002 USDC, which is a behavioral trait, but it does not state whether this is a read-only operation, what permissions are required, or any side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that efficiently conveys the core purpose and the cost. Every element earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity, the presence of an output schema, and the inclusion of cost, the description is largely complete. It lacks usage guidance and deeper behavioral disclosure, but for a straightforward classifier it is adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the parameters, and the description adds no additional meaning about 'action' or 'policy'. With high schema coverage, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Classify') and resource ('an action') with the three possible outcomes, making the tool's purpose clear. It does not explicitly distinguish itself from siblings like inter_agent_policy_evaluator or tool_call_policy_guard, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives. The description simply states what it does, without any context, prerequisites, or when-not-to-use conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
idempotency_replay_guardCInspect
Detect safe replays and conflicting idempotency-key reuse; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description carries the full behavioral burden. It discloses the cost (0.002 USDC) and that it detects replays/conflicts, but says nothing about side effects, authentication needs, rate limits, or what happens on detection. This is insufficient for a guard-type tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no wasted words. The price is tacked on after a semicolon, which is slightly awkward but still efficient and informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. However, with no annotations, the description should cover prerequisites, behavior, and usage context — it provides none of these, leaving significant gaps for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the two required parameters. The description adds no param-specific meaning (e.g., format of idempotency_key or allowed values for action), so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Detect') and resource ('safe replays and conflicting idempotency-key reuse'), making the tool's function clear. It does not explicitly differentiate itself from siblings like duplicate_charge_detector or concurrency_guard, but the idempotency-key focus is distinctive enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives, prerequisites, or exclusions. The description only states what it does and the cost; an agent must infer usage entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
index_feedPHION Index FeedBRead-onlyIdempotentInspect
Canonical catalog snapshot or digest-based delta with exact HTTP/MCP records, PHION Execute templates and signed evidence; 0.005 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | ||
| since_snapshot_sha256 | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful context such as the fee (0.005 USDC), the digest-based delta mechanism, and the exact record/template/evidence contents, but it does not explain response shape or pagination behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence that front-loads the core purpose and packs content, delta semantics, evidence, templates, and pricing without unnecessary filler. It is efficient, though the final clause is dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only, idempotent feed with two parameters, the description is adequate but not complete. It lacks explicit return-format details and does not fully connect the optional SHA-256 parameter to delta mode, which would matter given there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but the description partially compensates by naming 'snapshot' and 'digest-based delta,' which maps to the mode enum. It does not explicitly clarify when since_snapshot_sha256 should be supplied, though the pattern and parameter name make the intent inferable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource: a canonical catalog snapshot or digest-based delta, and enumerates its contents (HTTP/MCP records, PHION Execute templates, signed evidence). This goes beyond the name and distinguishes the tool from the many evidence/verification siblings, though it lacks an explicit verb.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'snapshot or digest-based delta' wording implies when the tool should be used, but it does not provide explicit when-to-use guidance, exclusions, or comparisons with sibling tools. Usage context is present mainly by inference from the content description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inspect_agentCRead-onlyIdempotentInspect
Pre-payment agent inspection across x402, A2A, ERC-8004, OpenAPI and MCP; failed inspections are not charged; 0.009 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and destructiveHint=false, but the description says successful inspections cost 0.009 USDC and failed inspections are not charged. Charging a fee is a side effect that contradicts the read-only annotation, so the description contradicts the structured behavioral hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with three semicolon-separated clauses; every clause adds distinct information: scope, failure billing, and price. There is no wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description does not explain what the inspection checks, what constitutes a failed inspection, or what result is returned. For a paid pre-payment inspection, this leaves the agent without enough context to interpret the outcome.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the lone 'url' parameter has no schema description. The protocol list weakly suggests the URL is an agent endpoint in one of those protocols, but it does not explain what URL should point to or what forms are accepted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: pre-payment agent inspection across named protocols. It is clear enough to distinguish from generic tools, though it does not explicitly differentiate itself from siblings such as payment_preflight or counterparty_risk_preflight.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Pre-payment' implies use before paying an agent, but there is no explicit when/when-not guidance or mention of alternatives such as payment_preflight, try_service, or verify. The agent must infer timing from the title-like clause rather than from stated usage rules.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inter_agent_policy_evaluatorInter-Agent Policy EvaluatorCRead-onlyIdempotentInspect
Require policy identity, subject, action, lifetime and valid decisions; compute the restrictive cap and shared actions; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| policies | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior. The description adds cost (0.003 USDC) and enumerates required fields, which is useful, but does not explain conflict behavior or how the computation is used. With annotations covering safety, this is an acceptable addition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very concise, but the three semicolon-separated clauses are poorly structured and front-loaded with a requirement rather than the tool's purpose. The information could be delivered more clearly in the same length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one complex array parameter, no output schema, and annotations that cover safety, the description is incomplete: it does not explain return values, how policies are evaluated, or the significance of 'restrictive cap' and 'shared actions'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must provide semantics. It lists the required fields ('policy identity, subject, action, lifetime and valid decisions') which maps to the schema's nested required fields, adding some meaning. However, it does not explain the array structure, min/max limits, or the meaning of each field beyond renaming.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific computation ('compute the restrictive cap and shared actions'), which is a clear verb+resource, but it is preceded by a requirement list and lacks a clear statement about evaluating policies or how it relates to sibling tools. It is not a tautology, but the purpose is not front-loaded.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like tool_call_policy_guard or preflight. The precondition list ('Require policy identity...') is a schema requirement, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
interrupted_task_recoveryCInspect
Resume only when task/checkpoint identity, sequence, deadline, attempt budget, idempotency and side effects are explicit; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | ||
| policy | Yes | ||
| checkpoint | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses a gating condition and a price of 0.003 USDC, but does not explain what happens if conditions are unmet, what the resume action actually mutates, or how idempotency and side effects are handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler, efficiently stating the gating condition and cost. The semicolon-separated price is compact, though the dense field list could be clearer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has three nested object parameters, 0% schema description coverage, no annotations, and no output schema. The description does not explain return behavior, failure modes, or how the nested objects map to recovery semantics, leaving significant gaps for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It names several required concepts (identity, sequence, deadline, attempt budget, side effects), but it omits nested fields such as checkpoint.state_sha256 and policy.max_attempts, and mentions idempotency which does not appear in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description implies a resume/recovery action for interrupted tasks, and the name confirms the resource. However, it does not state the purpose as a clear verb+resource pair and offers no differentiation from sibling recovery, idempotency, or transaction tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Resume only when...' gives a precondition for use, which is more than no guidance. But it does not say when to choose this tool over alternatives like transaction_recovery or idempotency_replay_guard, nor does it describe when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
journey_verifyBRead-onlyIdempotentInspect
Verify hash continuity, timestamps and ordered intent→quote→payment→delivery lifecycle; 0.010 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| events | Yes | ||
| expected | No | ||
| journey_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is lower. The description adds genuinely new behavioral context the annotations do not carry: what exactly is checked (hash continuity, timestamps, ordering) and the 0.010 USDC cost of the call. It still omits what a failure result looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the verification targets and tucks the price at the end. No filler, though the arrow notation for the lifecycle is compressed enough to require inference.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, no parameter documentation, and a nested 'expected' object, so the description should do more. It communicates scope and cost but leaves the return semantics (pass/fail shape, mismatch reporting) and the role of 'expected' unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across three parameters (journey_id, events, expected) with a nested object and array constraints (minItems 2, maxItems 128). The description vaguely gestures at event content through 'hash continuity' and 'ordered lifecycle' but never explains journey_id, what belongs in events, or the role of expected, so it does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (verify) and the exact resource scope: hash continuity, timestamps, and an ordered intent→quote→payment→delivery lifecycle. That is far more concrete than the sibling 'verify' or 'transaction_assurance', though it does not explicitly name how it differs from those near-neighbors.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the lifecycle wording — you call this to validate a recorded journey's integrity — but there is no explicit when-to-use statement, no prerequisites, and no pointer to alternatives like payment_receipt_reconciler or transaction_assurance for overlapping verification needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
live_data_freshnessCInspect
Evaluate live source observation against explicit maximum age; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| source_urls | Yes | ||
| max_age_seconds | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It only adds a cost figure (0.002 USDC) but does not disclose the return value, how the evaluation works, what happens when the maximum age is exceeded, or any side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very brief and front-loads the action and cost, but it is a fragment rather than a structured explanation. The brevity leaves out necessary details, making it under-specified rather than efficiently concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description is far from complete. An agent would not know what result to expect, how to handle stale sources, or what the 0.002 USDC signifies beyond a cost.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the parameters. It implicitly maps 'explicit maximum age' to max_age_seconds but does not clarify source_urls beyond a vague 'live source observation', leaving the URL format and interpretation undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Evaluate') and resource ('live source observation against explicit maximum age'), indicating a freshness check. It is not a tautology and is distinct enough from related tools, though it doesn't explicitly compare to siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives. The description only includes the task statement and a cost, with no mention of context, exclusions, or sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mandate_reserveCIdempotentInspect
Atomic spending reservation with cumulative cap and conflict-safe idempotency; 0.004 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| amount | Yes | ||
| mandate_id | Yes | ||
| ttl_seconds | No | ||
| idempotency_key | Yes | ||
| capability_token | Yes | ||
| cumulative_limit | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true, readOnlyHint=false, and destructiveHint=false, so the safety profile is covered. The description adds genuine context beyond that: atomicity, a cumulative cap that persists across calls, and conflict-safe idempotency. It still omits what happens on cap exhaustion, whether a reservation can be released/expires (ttl_seconds), and the role of capability_token as an authorization requirement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the key traits (atomic, capped, idempotent) come first and the cost note is appended. It is dense but not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a state-mutating, 6-parameter, 5-required financial tool with no output schema and zero schema descriptions, and the description does not compensate. Authorization via capability_token, TTL behavior, cap-exhaustion behavior, and return semantics are all unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 6 parameters, so the description carries the full burden and largely fails. 'Cumulative cap' loosely maps to cumulative_limit and 'idempotency' to idempotency_key, but amount units (the schema only gives an integer-string pattern), capability_token, ttl_seconds, and mandate_id are never explained. The '0.004 USDC' figure describes the tool's own price, not the amount parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action and object: an atomic spending reservation against a mandate, with a cumulative cap and idempotency guarantee, plus its price (0.004 USDC). The verb+resource is clear, but it does not differentiate itself from plausible siblings such as agent_budget_guard, subscription_spend_guard, idempotency_replay_guard, or rwa_nav_reserve, leaving the agent to guess which reserve/guard tool applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never says when to call this versus the many guard/preflight siblings, nor does it state prerequisites such as holding a valid mandate_id or capability_token. Usage is only implied by the phrase 'spending reservation'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
manifest_version_diffBInspect
Recursive manifest diff that requires explicit versions and flags sensitive changes without a version advance; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| current | Yes | ||
| previous | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that the diff is recursive, requires explicit versions, flags sensitive changes without a version advance, and costs 0.002 USDC. However, it does not explain what counts as sensitive, how the diff is returned, or any permission or payment workflow beyond the cost mention.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It efficiently combines the core operation, a key constraint, a behavioral flag, and pricing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with nested inputs, no output schema, and no annotations, the description is moderately complete: it conveys the diff behavior, version requirement, sensitive-change flagging, and cost. Still missing are return format, treatment of nested manifest fields, and any auth or payment details beyond the stated USDC amount.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and both parameters are nested objects. The description only adds that explicit versions are required, which maps to the schema's required version fields, but it does little to explain the remaining object structure or how previous and current manifests are compared. With low coverage and nested objects, this is insufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific operation (recursive manifest diff) and adds distinctive conditions: explicit versions required and sensitive changes flagged only when there is no version advance. It is clear what the tool does, though it does not explicitly distinguish itself from siblings like mcp_manifest_firewall or tool_capability_drift_monitor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, nor when not to use it. The description implies a diff use case, but it leaves the agent to infer the correct context and any prerequisites from the name and function alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
market_data_snapshotCInspect
Timestamped public market-data snapshot with provenance; 0.005 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| subject | No | ||
| source_urls | Yes | ||
| expected_fields | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It reveals that the snapshot is timestamped, provenance-bearing, and costs 0.005 USDC, which is useful context. But it does not disclose side effects, data-source behavior, failure modes, or what the returned snapshot structure is, leaving major behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the core concept and ends with the cost. It contains no filler words. It loses one point because it prioritizes compactness over including a verb and parameter linkage, but structurally it is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 3 parameters, 0% schema coverage, no annotations, and no output schema, the description needs to carry substantial weight. It does not explain how to call the tool, what the output looks like, or what the required parameters mean. The brief noun phrase and price are insufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining the three parameters. Instead, it only loosely maps to the idea of market data and provenance and never mentions source_urls, subject, or expected_fields. The required source_urls parameter is completely unexplained, leaving an agent without enough information to fill it correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description says "Timestamped public market-data snapshot with provenance; 0.005 USDC." This names a specific resource and adds distinguishing attributes (timestamped, provenance, cost), which separates it from many siblings. However, it lacks a verb — it never states explicitly what the tool does with the snapshot — so it stops just short of a full purpose statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives. It implies a paid public-data retrieval context with the cost note, but it does not state when to prefer this over sibling tools like live_data_freshness or source_backed_search, nor any conditions or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
market_demand_predictorMarket Demand PredictorARead-onlyIdempotentInspect
Forecast verified demand trends while abstaining on insufficient history.; 0.005 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, and non-destructive behavior. The description adds valuable traits beyond annotations: determinism, evidence-boundedness, signing, and the abstaining behavior on insufficient history. This provides meaningful context for an agent's expectations, though it does not detail output format or error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and front-loads the core purpose. Each phrase adds value (purpose, cost, behavioral traits). However, the punctuation is awkward (multiple periods/semicolon) and the cost mention could be considered extraneous, but overall it is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with an open input schema, no parameter guidance, and only an output schema (not described), the description is incomplete. An agent cannot determine what input to provide or what the response structure will be. The abstaining behavior is noted, but the lack of input/output details makes the tool hard to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero defined parameters but is an open object requiring at least one property (minProperties: 1). The description does not explain what properties to provide, leaving the agent without guidance on input structure. With no named parameters, the baseline is 4, but the open schema creates a gap that the description fails to fill.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Forecast verified demand trends') and a resource ('market demand'), and includes a behavioral qualifier ('abstaining on insufficient history') that differentiates it from generic predictors. The added traits (deterministic, evidence-bounded, signed) further clarify its distinct nature among many sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a usage condition ('abstaining on insufficient history') but does not explicitly name alternatives or state when this tool should be preferred over other predictors like provider_quality_predictor or price_discovery_engine. There is no explicit 'use this when' guidance or exclusion of other contexts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcp_2026_compatibility_gatewayMCP 2026 Compatibility GatewayCRead-onlyIdempotentInspect
MCP 2026 Compatibility Gateway; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly, idempotent, and non-destructive behavior. The description adds 'deterministic' and 'signed' attributes, which slightly extends the behavioral picture by indicating reproducible, cryptographically signed results. However, it does not disclose what input boundaries exist or what 'assessment' actually entails beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one sentence and wastes no words. Front-loading the name and appending the key behavioral qualifier is efficient. It is concise, though the brevity contributes to vagueness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With two required parameters and no enums, the tool is moderately simple, but the description leaves essential details undefined: what protocol versions are valid, how methods should be passed, what a 'signed assessment' looks like, and how the result should be consumed. The output schema mitigates return-value ambiguity, but input semantics and operational meaning are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema names required parameters (protocol_version, methods) but gives no property descriptions Armstrong. The description's 'supplied input' adds no meaning to these parameters. Although schema description coverage is reported as 100%, the visible schema lacks any property-level documentation, so the description does not compensate for the missing semantic context around protocol_version or methods.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description, 'deterministic signed assessment over supplied input,' restates the name's 'compatibility gateway' without explaining what compatibility is being assessed or what the output represents. There is no clear verb-resource pair—'assessment' is generic and does not distinguish this from numerous sibling assessment/gateway tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives. None of the sibling tools (e.g., mcp_manifest_firewall, preflight, verify) are referenced, and no context or exclusions are provided, leaving an agent with no basis for routing between similar tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcp_a2a_task_bridgeMCP-A2A Task BridgeCRead-onlyIdempotentInspect
MCP-A2A Task Bridge; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, and the description adds 'deterministic signed assessment'—behavioral traits not present in the annotations. This provides meaningful extra context without contradicting the declared hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler words. It is concise, though the brevity arguably contributes to the vagueness. Still, it earns its place as a quick tagline.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description fails to explain what mcp_task and a2a_task are or what form they should take, leaving the agent unable to construct a valid invocation. The output schema exists, but the input side is completely undefined both in the schema and in the description, making the tool incomplete for practical use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has no property definitions and parameter count is 0, so the baseline is 4. The description does not add any semantic meaning to the required fields mcp_task and a2a_task, but since the schema itself has no descriptions, the description is not expected to compensate beyond the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description says 'MCP-A2A Task Bridge; deterministic signed assessment over supplied input' but does not state what the tool actually does with the input or what kind of assessment it produces. The verb/resource ('assessment over supplied input') is vague and does not clearly distinguish it from siblings like a2a_transaction_bridge or ap2_mandate_bridge.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no guidance on when to use this tool, when not to use it, or how it compares to alternatives. There is no mention of prerequisites, selection criteria, or exclusions, leaving the agent to guess based on the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcp_catalog_cache_guardMCP Catalog Cache GuardDRead-onlyIdempotentInspect
MCP Catalog Cache Guard; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the notion of 'deterministic signed assessment' hinting at signature generation but provides no detail on what is signed, what the output contains, or any side effects. This is minimal extra context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely short (one sentence) but it is under-specified rather than concise. It fails to convey essential information, making it ineffective. The brevity is not beneficial because it sacrifices clarity; it does not 'earn its place' by providing any actionable detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three required parameters and a rich set of siblings, this description is wholly inadequate. It does not explain what the assessment does, what the parameters mean, or what the output schema provides. Even with annotations and an output schema, an agent cannot determine when to call this tool or what inputs to supply. The description is far from complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is reported at 100%, implying each parameter (catalog, ttl_ms, cache_scope) has a description in the actual schema. Since the schema already documents these, the description need not repeat them. However, the description adds no additional semantic nuance beyond 'over supplied input', so it earns the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'deterministic signed assessment over supplied input' which gives a vague sense of purpose (assessment) but does not specify the resource, action, or scope. It fails to differentiate from numerous sibling tools with similar names like 'cache' related or 'guard' tools, and the purpose is ambiguous. It is not a tautology but is far from clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is absolutely no guidance on when to use this tool versus any of the many siblings. No context, no alternatives, no conditions for invocation. The description is silent on usage scenarios entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcp_manifest_firewallCInspect
Inspect an MCP manifest before installation or trust; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| policy | No | ||
| manifest | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses a cost (0.002 USDC), which is a real trait, but says nothing about what the inspection detects, whether it blocks installation, what happens on a failing manifest, or any auth requirements. The safety profile is essentially unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence plus a price tag; no filler. It may be too terse for a nested-object tool, but it wastes nothing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool takes nested objects, has zero schema description coverage, and no annotations, yet the description omits the policy parameter, the detection semantics, and the meaning of a pass/fail result. An output schema exists so return values needn't be explained, but the gaps in scope and parameters leave it inadequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% with two nested-object parameters, so the description must compensate and does not. The required 'manifest' object is somewhat inferable, but the optional 'policy' object is completely opaque in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Inspect) and resource (MCP manifest) with a clear timing qualifier (before installation or trust). It distinguishes itself reasonably from siblings like tool_output_firewall, though the exact scope of 'inspect' is left fuzzy.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'before installation or trust' gives an implied trigger condition, so an agent knows roughly when to reach for it. However, it names no alternatives and offers no exclusions, which is weak given the many adjacent preflight/guard siblings (preflight, tool_call_policy_guard, counterparty_risk_preflight).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcp_server_identity_evidenceCInspect
Bind MCP domain, endpoint, manifest, key and version; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It reveals the monetary cost (0.003 USDC), which is useful, but says nothing about whether the operation is read-only or mutating, what permissions are required, or what side effects binding may have.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It efficiently states the action, the components involved, and the cost, though it may be too terse for a tool with financial and evidence implications.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values do not need explanation. However, with no annotations and a complex, payment-bearing identity-binding operation, the description omits critical context such as when to invoke it, what it produces, and how it relates to sibling evidence tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the required 'identity' parameter. The description adds a list of components (domain, endpoint, manifest, key, version) that likely map to fields within that identity, but does not clarify the parameter's structure or format beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb, 'Bind', and enumerates the identity components it operates on: MCP domain, endpoint, manifest, key, and version. It also names the cost, which helps identify this as a paid evidence-binding operation. However, it does not distinguish itself from adjacent siblings such as domain_ownership_evidence or mcp_manifest_firewall.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives, nor does it state preconditions, exclusions, or typical workflows. The only contextual clue is the price, which does not explain usage in relation to sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mcp_transaction_gatewayMCP Transaction GatewayARead-onlyIdempotentInspect
Bind an MCP tool call to transaction evidence without granting undeclared authority.; 0.004 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, which paint a safe, non-mutating picture. The description adds meaningful behavioral context by specifying that authority is 'not undeclared', that it is 'deterministic' and 'evidence-bounded', and that it produces a signed output. This clarifies the tool's safety and output nature beyond the annotations. Note: the 'without granting undeclared authority' phrase is a bit ambiguous but does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: one sentence plus a small list of attributes. It front-loads the primary purpose and then lists key properties without any fluff. Every word contributes to understanding the tool's behavior and limitations. This is an example of minimal length with high information density.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema (not shown but present) and this is a special-purpose security tool, the description covers the core purpose and guarantees. However, it leaves out important context: it doesn't explain what constitutes 'transaction evidence', how the input should be structured (given the open schema), or how to interpret the signed result. An agent might be uncertain about the exact contract for invocation. The description is adequate but not fully complete for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters in the description (though the input schema allows additional properties, meaning it accepts arbitrary input). With no parameters to document, the description has little to add on parameter semantics. The description mentions 'Bind an MCP tool call' which implies there is an input referencing the tool call, but it doesn't explain the expected structure of that input. Given the open schema and zero parameter coverage, the description could elaborate on what the input should contain, but it remains at a baseline level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the core function: binding an MCP tool call to transaction evidence, with specific characteristics (deterministic, evidence-bounded, signed). It goes beyond a simple verb+resource by highlighting the security guarantees. However, it does not clearly distinguish itself from the many sibling tools like transaction_assurance, payment_preflight, or mcp_2026_compatibility_gateway, which could plausibly overlap in purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for when you need to bind a tool call to transaction evidence, but it does not explicitly state when to use this tool versus alternatives (e.g., transaction_assurance, proof_of_service). It lacks explicit exclusion criteria or references to sibling tools, so an agent might not know whether to choose this or another evidence-related tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_write_guardBInspect
Screen persistent memory writes for poisoning, unsafe instructions and missing provenance; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| write | Yes | ||
| policy | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It does add one genuinely useful behavioral fact — the 0.002 USDC cost — and 'screen' implies a non-destructive check, but it omits whether the call blocks the write, what happens on a failed screen, and any auth/payment requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded clause stating action, target, and criteria, with the cost appended. Every element earns its place and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, but with zero annotation coverage, two undocumented nested-object parameters, and no guidance on failure handling or policy semantics, the description is too thin for a guard tool that gates writes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two nested-object parameters, and the description only loosely gestures at the 'write' parameter via the phrase 'persistent memory writes'. The 'policy' object is never mentioned, so the description does not compensate for the documentation gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Screen') plus resource ('persistent memory writes') and enumerates the three checks performed (poisoning, unsafe instructions, missing provenance). It is clear what the tool does, though it offers no explicit contrast against potentially confusable siblings like tool_output_firewall or schema_normalize.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to invoke this guard versus alternatives, no prerequisites, and no exclusions. The price suffix hints this is a paid service (siblings include a *_free variant elsewhere), but the description never tells the agent to prefer a cheaper path or when screening is mandatory.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
multi_agent_escrowMulti-Agent EscrowARead-onlyIdempotentInspect
Recommend hold or release from evidence without PHION custody or fund movement.; 0.005 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior, and the description adds non-obvious traits: an exact 0.005 USDC cost, deterministic output, evidence-bounded reasoning, and signed results. These details materially shape an agent's expectations and are not visible in the schema or annotations. No contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the action, then adds cost and behavioral properties in separate clauses. Every fragment contributes distinct information and there is no filler, though the telegraphic style is slightly terse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with rich annotations and an output schema, the description covers purpose, cost, determinism, safety boundaries, and a signed result. The only mild gap is that the open input schema is not described, but 'from evidence' gives enough of a hint for an agent to pass evidence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero declared parameters and schema description coverage is 100%, so the schema carries no parameter burden. With no parameters, a baseline score of 4 is appropriate; the description does not need to elaborate on inputs, though it leaves the open additionalProperties object unspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific action — 'Recommend hold or release' — and scopes it to evidence, while explicitly excluding custody and fund movement. It is clear enough to distinguish the tool from custody/movement siblings, though it does not name alternatives the way the strongest definitions do.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use statement, required precondition, or named alternative, so an agent must infer when to select this tool over the many related decision, escrow, and payment siblings. The 'without PHION custody or fund movement' clause only states a boundary, not selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
multi_source_fact_bundleCInspect
Independent source observations with agreement made explicit; 0.005 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| source_urls | Yes | ||
| expected_fields | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of disclosing effects, authentication, or data handling, but it only adds the price and a rough characterization of the result. It does not state whether this creates records, performs verification, or requires certain input preparation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The text is brief, front-loaded, and free of filler; the cost is stated compactly. However, the economy is achieved by omitting substantive guidance, so it is concise but not sufficiently informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has three parameters, including a required source_urls array, no output schema, and no annotations, yet the description explains none of the invocation semantics or expected return shape. Critical information about how the multi-source agreement is constructed and what the agent receives is entirely absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention query, source_urls, or expected_fields at all. An agent must infer parameter meaning solely from property names and types, so the description completely fails to compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a noun phrase describing what the bundle contains, not an action an agent can invoke. 'Independent source observations with agreement made explicit' essentially restates the tool's name without a verb like 'retrieve' or 'aggregate,' and it doesn't distinguish this tool from siblings such as source_backed_search or fetch_evidence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool over sibling evidence tools, nor any prerequisite contexts or exclusions. The only additional detail is the price, which is an economic constraint, not a usage criterion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oauth_issuer_binding_evidenceOAuth Issuer Binding EvidenceCRead-onlyIdempotentInspect
OAuth Issuer Binding Evidence; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the behavioral traits 'deterministic' and 'signed', which are useful, though 'deterministic' largely overlaps with the idempotency annotation. It does not describe authentication needs, rate limits, or failure behavior, but it does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no filler or redundancy beyond echoing the title. The phrase 'deterministic signed assessment over supplied input' is front-loaded and easy to scan. It is concise, though its brevity limits substance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool involving OAuth issuer binding and signed evidence, the description leaves too much unexplained: what 'issuer binding' means, what the assessment output looks like, and how the required parameters interact. Output schema and annotations cover return shape and safety, but the description itself is not complete enough for an agent to confidently know when or why to invoke this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is reported as 100%, so the parameters claims, expected_issuer, and expected_audience are already documented in the schema. The description itself adds no parameter-level explanation, which is acceptable under the high-coverage baseline, but it also does not enrich the meaning of those parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states that this tool produces a 'deterministic signed assessment over supplied input', which hints at an assessment/evidence-generation operation. However, it never says what the assessment verifies about OAuth issuer binding or how it relates to the required claims, expected_issuer, and expected_audience parameters. It is more specific than a pure tautology but remains vague and does not distinguish it from sibling evidence tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as oauth_token_audience_guard, mcp_server_identity_evidence, or other evidence-generation tools. The description provides no prerequisites, exclusions, or typical use cases, leaving selection entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
oauth_token_audience_guardBInspect
Validate declared OAuth audience and issuer; never send raw tokens; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses that raw tokens are never sent and mentions a cost of 0.002 USDC, but it does not describe failure behavior, permissions, side effects, or limits of the validation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with three distinct clauses, front-loading the core validation purpose and adding only high-value details about token handling and cost.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. However, with no annotations, the description should do more to clarify usage, error handling, or the validation outcome for an agent to invoke it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is reported as 100%, so the baseline is 3. The description adds no parameter-level meaning beyond what the schema already provides for claims and expected_audience.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: validate OAuth audience and issuer. It is clear what the tool does, but it does not explicitly distinguish itself from the many sibling guard/preflight tools, nor does it name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, nor any when-not conditions. The description implies a validation checkpoint but leaves usage context to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onchain_evidenceCInspect
Public explorer or RPC response evidence with immutable hashes; 0.004 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| subject | No | ||
| source_urls | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears the full burden of behavioral disclosure. It mentions immutable hashes and a cost of 0.004 USDC, which is useful, but it does not state side effects such as whether the tool performs network fetches, writes evidence, verifies hashes, or what happens on failure. This is minimal transparency for an on-chain evidence operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, with no filler words. However, it is terse to the point of ambiguity, and the cost annotation is placed awkwardly after a semicolon rather than integrated into a clear behavioral explanation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With two parameters, no output schema, and no annotations, this single cryptic sentence is completely inadequate for an agent to select and invoke the tool correctly. It omits the action performed, the meaning of subject, how source_urls are used, what the evidence output looks like, and when this tool should be preferred over any sibling evidence tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for both source_urls and subject. The phrase 'Public explorer or RPC response evidence' hints that source_urls should point to public explorer or RPC endpoints, but it never mentions the subject parameter or explains its role. The description adds only weak contextual meaning beyond the schema's bare parameter list.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource as 'Public explorer or RPC response evidence with immutable hashes', which gives some clue that this tool deals with on-chain evidence and hashes. However, it lacks an explicit verb or clear action statement, so an agent cannot tell whether this tool creates, verifies, or fetches evidence. It also does not clearly distinguish itself from the many sibling evidence tools beyond the 'public explorer or RPC' phrase.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance about when to use this tool versus siblings such as fetch_evidence, delivery_evidence, or domain_ownership_evidence. The phrase 'Public explorer or RPC response evidence' implies a use case, but no conditions, exclusions, or alternative recommendations are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
payment_capability_preflightDeclare Payment CapabilityBRead-onlyIdempotentInspect
Free anonymous non-blocking self-report of wallet, mandate, x402 v2, network, asset, account kind and funding readiness. Detects EVM account types incompatible with exact EIP-3009 before signature. Stores only a daily pseudonymous actor and normalized reason; never a wallet, balance, signature or request content.
| Name | Required | Description | Default |
|---|---|---|---|
| intent | Yes | ||
| service | Yes | ||
| funds_available | No | ||
| evm_account_kind | No | ||
| spend_authorized | No | ||
| supported_assets | No | ||
| supports_x402_v2 | No | ||
| wallet_available | No | ||
| supported_networks | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent and non-destructive, so the baseline is covered. Beyond that, the description discloses real behavioral traits: the call is non-blocking, it stores only a daily pseudonymous actor and normalized reason, and it never stores a wallet, balance, signature or request content, plus it detects incompatible EVM account types before signature. This privacy and detection detail materially exceeds what the annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with what the tool reports and then the safety/detection and data-handling guarantees. Dense but each clause carries distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should ideally say what the caller receives back from this detection/report, and it does not. It is complete on privacy and behavior but silent on return semantics and on usage relative to sibling preflight tools, leaving an agent with reasonable but not full context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning, and it does enumerate the readiness dimensions (wallet, mandate, x402 v2, network, asset, account kind, funding) that map onto the nine inputs. However, it gives no format or value guidance, and the required 'service' and 'intent' parameters are never explained, leaving gaps for a 9-parameter schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource: a self-report of wallet, mandate, x402 v2, network, asset, account kind and funding readiness, plus an EIP-3009 compatibility check. This is far more concrete than a tautology. It does not, however, distinguish itself from close siblings such as payment_preflight or preflight, so sibling differentiation is missing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says the call is 'free anonymous non-blocking,' which hints it is cheap and safe to invoke, but it never states when to use this instead of payment_preflight, preflight, or payment_diagnose. No prerequisites, no exclusions, no routing guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
payment_delivery_atomicityPayment Delivery AtomicityDRead-onlyIdempotentInspect
Payment Delivery Atomicity; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the description need not repeat those safety signals. However, the description adds no behavioral context beyond the annotations: it does not disclose what happens to the input, what kind of assessment is produced, whether it returns evidence, or any side effects. It does not contradict annotations, but it provides zero added value, and the description is so thin that it fails to explain the tool's actual behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely short, one clause plus a phrase. While conciseness is generally good, this is under-specification rather than effective conciseness. The single clause restates the name and adds minimal value. It is not front-loaded with useful information because there is essentially no useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema, the tool's purpose is opaque. The description gives no sense of what the atomicity assessment entails, how it relates to payment and delivery, or what the output schema represents. For a tool named 'Payment Delivery Atomicity' with two required parameters and a clear domain, a complete description needs at least a statement of what atomicity means in this context and what is assessed. This is completely inadequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has two required parameters, 'payment' and 'delivery', and has 100% schema description coverage (via their names). The description does not add any explanation of these parameters beyond what the names imply. Since schema coverage is high, the baseline is 3, and the description does not exceed that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is essentially a tautology of the name/title: 'Payment Delivery Atomicity; deterministic signed assessment over supplied input.' It does not state what the tool does—only that it produces a deterministic signed assessment over some input. It does not mention payment or delivery atomicity, which are the core concepts in the name. The description is vague and fails to distinguish this tool from the many other assessment/verification tools in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as payment_preflight, payment_receipt_reconciler, or transaction_assurance. The description provides no context about the problem it solves, no prerequisites, and no exclusions. It simply restates the tool's existence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
payment_diagnoseCInspect
Free diagnosis of the exact payment-funnel stage and machine-readable recovery actions for any PHION service.
| Name | Required | Description | Default |
|---|---|---|---|
| service | Yes | ||
| body_valid | No | ||
| error_code | No | ||
| http_status | No | ||
| x402_version | No | ||
| selected_network | No | ||
| has_payment_signature | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose two real traits beyond the schema: the call is free (a cost trait) and it returns machine-readable recovery actions. However, it omits whether the operation is read-only, what auth or rate limits apply, and what the recovery output looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-formed sentence with the key distinction (exact stage + recovery actions) front-loaded and no filler. It is efficient, though arguably too terse given the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 7 parameters, no annotations, and no output schema, one sentence leaves major gaps: parameter formats and interrelationships (body_valid, error_code, http_status) and any return-shape guidance are absent. A caller has only the bare purpose to work from.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Seven parameters with 0% schema description coverage and no enums, yet the description explains none of them. Concepts like service, error_code, x402_version, or has_payment_signature and their interplay are left entirely unaddressed, so it fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
It names a specific verb and resource ('diagnosis of the exact payment-funnel stage') plus the output ('machine-readable recovery actions'), which is far more informative than a tautology. It is not fully disambiguated from close siblings like payment_preflight or transaction_recovery, but an agent can reasonably tell it apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The only usage cue is the scope 'for any PHION service' and the word 'Free'. There is no statement of when to prefer this over payment_preflight, transaction_recovery, or try_service, and no prerequisites are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
payment_optimizerPayment OptimizerBRead-onlyIdempotentInspect
Rank only integration-tested payment routes under explicit cost and latency constraints.; 0.003 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, and the description adds useful behavioral traits: 'Deterministic' reinforces idempotency, 'evidence-bounded' indicates inputs are constrained to verified data, and 'signed' suggests the output has a cryptographic signature. These go beyond the structured hints without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is highly compact and front-loaded with the core action. However, the fragment '; 0.003 USDC.' is syntactically awkward and unexplained, and it consumes a sentence slot without clear value. The remaining phrases are informative but the overall structure feels clipped rather than intentionally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with an open input schema and no parameters, the description leaves too much unsaid: how constraints are provided, what 'integration-tested' means in practice, what the signed output contains, and how this differs from payment_route_selector. The output schema exists but is not shown in context, so the description alone does not give an agent enough to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With a parameter count of 0 and 100% schema description coverage, the baseline is 4. The description mentions 'explicit cost and latency constraints' but does not specify how these are expressed (likely because the schema is an open object). Given no formal parameters, this is acceptable, though a little more context on the expected constraint shape would be helpful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Rank'), a specific resource ('payment routes'), and a scope ('only integration-tested', 'under explicit cost and latency constraints'). It does not explicitly distinguish itself from sibling tools like payment_route_selector, but the scoping language is specific enough that an agent can infer its niche.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, no exclusions, and no mention of prerequisites. The phrase 'only integration-tested' implies a filtering behavior but does not explain how to decide between payment_optimizer and payment_route_selector or other payment-related siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
payment_preflightCRead-onlyIdempotentInspect
Fail-closed x402 v2 firewall for network, asset, recipient, amount, scheme, resource host and expiry; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| policy | Yes | ||
| intent_id | No | ||
| payment_requirement | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds one genuinely useful behavioral trait — 'fail-closed' (denies by default) — plus a cost signal ('0.002 USDC'), but says nothing about required permissions, what a denial returns, or rate limits. Because annotations carry most of the burden, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Effectively a single front-loaded clause with no filler. The trailing cost figure sits awkwardly and would fit better as its own sentence, but the definition is not bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, but the tool has nested required objects with 0% schema coverage and the description never defines the shape or accepted values of 'policy'/'payment_requirement'. For a fail-closed gate whose correctness hinges on those inputs, this leaves too much for the agent to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and two of three parameters (policy, payment_requirement) are nested objects. The enumerated check dimensions loosely sketch what payment_requirement contains, but the required 'policy' object and 'intent_id' are entirely unexplained, so the description does not compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific screening gate ('fail-closed x402 v2 firewall') and enumerates the dimensions it checks (network, asset, recipient, amount, scheme, resource host, expiry), which is far more specific than a tautology and helps distinguish it from sibling tools like x402_quote_comparator or payment_diagnose. It is somewhat jargon-dense, but an agent can tell what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use, when-not, or routing to alternatives (payment_diagnose, payment_route_selector, x402_quote_comparator, counterparty_risk_preflight are all plausible siblings). The prefix 'preflight' weakly implies 'before sending a payment,' but no condition or prerequisite is stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
payment_receipt_reconcilerPayment Receipt ReconcilerCRead-onlyIdempotentInspect
Compare canonical and common x402 payment/receipt aliases, settlement state and coverage; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| payment | Yes | ||
| receipt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description usefully adds the cost signal (0.003 USDC), which is real behavioral context not present in annotations. It says nothing, though, about what the comparison produces or how mismatches are reported, so the added value is limited to the price.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the verb and objects first and the price appended; no filler. It is efficient, though the density of undefined jargon limits how much it actually communicates.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Two required nested-object parameters with zero schema documentation and no output schema leave the agent without enough information to construct a correct call. For a reconciliation tool that is expected to compare structured payment and receipt data, the definition is under-specified about both input shape and expected result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and both required parameters are untyped nested objects, so the schema gives the agent no field-level meaning. The description references 'payment/receipt aliases' and 'settlement state', which gestures at expected content but does not specify any accepted keys or structure for the two objects.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a verb (compare) and the two resources (payment, receipt), plus mentions settlement state and coverage, so the general intent is graspable. However, terms like 'aliases' and 'coverage' are unexplained jargon, and nothing distinguishes it from siblings such as payment_diagnose, duplicate_charge_detector, or transaction_assurance. Purpose is implied rather than crisply defined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus the many adjacent payment/x402 tools in the sibling list (payment_diagnose, payment_preflight, x402_quote_comparator). Neither preconditions nor exclusions are stated; the agent is left to infer usage entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
payment_route_selectorPayment Route SelectorARead-onlyIdempotentInspect
Reject impossible x402 routes, rank safe eligible routes by policy and return the exact next payment action; read-only, deterministic, 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| policy | Yes | ||
| routes | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/non-destructive/openWorld=false, so the bar is lower. The description adds that processing is deterministic and carries a 0.002 USDC cost, which is genuinely useful pricing context not present in structured fields. It doesn't say what 'reject impossible' entails or what the output contains, but output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence that front-loads the reject/rank/return pipeline and ends with the safety+cost qualifiers. No waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values needn't be explained. However, for a tool with deeply nested policy and route structures at 0% param description coverage, the description leaves the agent guessing about semantics of the filter fields, which is a meaningful gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across many nested fields (policy.* and routes.*). The description mentions 'routes' and 'policy' by name only, adding no meaning about allowed_assets, preferred_networks, max_total_atomic, or any route field. With low coverage, the description should compensate and does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action chain: reject impossible routes, rank eligible ones, return the next action. Distinguishes itself as a selector/ranker rather than a comparator or diagnoser among siblings. Does not name which sibling it replaces, so short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implied usage: use when you have candidate x402 routes plus a policy and need the next payment action. No explicit when-to-use vs payment_preflight, payment_diagnose, or x402_quote_comparator despite this being a crowded payment-tool family.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
person_enrichment_evidenceCInspect
Person enrichment with free input preflight, provider attribution, likelihood, limitations and signed receipt; 0.006 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| lid | No | ||
| name | No | ||
| No | |||
| input | No | ||
| phone | No | ||
| query | No | ||
| entity | No | ||
| pdl_id | No | ||
| person | No | ||
| region | No | ||
| school | No | ||
| ticker | No | ||
| company | No | Company name or nested input object | |
| contact | No | ||
| country | No | ||
| profile | No | ||
| subject | No | ||
| website | No | ||
| locality | No | ||
| location | No | ||
| last_name | No | ||
| birth_date | No | ||
| email_hash | No | ||
| first_name | No | ||
| identifier | No | Email, profile URL, phone, company domain, ticker or name; interpreted by service | |
| postal_code | No | ||
| street_address | No | ||
| minimum_likelihood | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral transparency burden. It does add useful details: free input preflight, provider attribution, likelihood, limitations, signed receipt, and a price of 0.006 USDC. However, it does not disclose authentication needs, side effects, rate limits, or what the returned evidence/receipt structurally contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler words, and it front-loads the core purpose before listing notable features and cost. It is appropriately brief, though the semicolon-joined feature list reads a bit densely.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 28 parameters, no output schema, and no annotations, so the description is severely insufficient. It does not explain how to construct a valid request, what fields are meaningful together, what the signed receipt looks like, or how preflight and validation behave. An agent would be guessing on most operational details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 7% across 28 parameters, and the tool description does not compensate. It mentions 'person' and 'free input' but does not explain any of the parameters, their roles, required combinations, or how the service interprets them. The description adds almost no semantic value over the sparse schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific action and resource: 'Person enrichment'. It also distinguishes itself from sibling enrichment tools like company_enrichment_evidence and contact_enrichment_evidence by naming 'person' as the subject. However, the phrase is somewhat jargon-heavy and doesn't fully explain what 'enrichment evidence' means in practice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus alternatives. It does not mention exclusions, prerequisites, or conditions that would route an agent to a different sibling tool. The only implicit hint is the word 'person', which is not enough for a tool with so many siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phion_discoverDiscover PHION ServicesARead-onlyIdempotentInspect
Free complete PHION service index generated from canonical inventory. Use when the agent knows a capability or wants to inspect the full network.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, covering the safety profile. The description adds that the index is 'free' and 'generated from canonical inventory', which is useful freshness/provenance context, but it says nothing about pagination, size, or format of the full index.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loading the nature of the resource before the usage clause. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-param read tool with full annotation coverage, the safety and usage essentials are present, but the 'complete index' likely returns a large payload and the description doesn't hint at size or output shape (no output schema exists to compensate). Could say more about what the index contains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description correctly implies no filtering input is required, consistent with a full-index operation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: provides a 'complete PHION service index'. It says what the tool returns and frames it as an inventory inspection. It doesn't explicitly distinguish itself from sibling phion_search_services, which likely offers a filtered view.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use when the agent knows a capability or wants to inspect the full network' gives a usage condition, but it doesn't contrast against siblings like phion_search_services or phion_get_service. The guidance is implied rather than explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phion_enrichPHION EnrichARead-onlyIdempotentInspect
Enrich a company, person or contact through an attributed provider route without hiding upstream cost.; 0.120 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnly, idempotent, and non-destructive behavior. The description adds genuinely useful behavioral traits: 'Deterministic, evidence-bounded and signed,' plus the exact cost and a statement about not hiding upstream cost. No annotation contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loads the core action and resource. The punctuation is slightly awkward ('without hiding upstream cost.; 0.120 USDC.'), but overall every fragment contributes meaningful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a simple read-only, idempotent enrichment call, and an output schema exists. However, given the large sibling tool list and an open input schema, the description does not fully clarify how to structure the request or how this tool relates to the enrichment evidence tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0 formally defined parameters and an open-envelope schema, the baseline is 4. The description provides target-type semantics ('company, person or contact'), which is helpful, though it still does not specify what fields or identifiers a caller should supply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Enrich') on a specific resource type ('a company, person or contact') and adds the 'attributed provider route' qualifier. However, 'enrich' is not formally defined, and the description does not explicitly distinguish phion_enrich from phion_search, phion_inference, or the separate enrichment evidence siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given for when to use this tool versus sibling tools such as phion_search, phion_inference, company_enrichment_evidence, person_enrichment_evidence, or contact_enrichment_evidence. There are no when-to-use, when-not-to-use, or alternative conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phion_enrich_basicPHION Basic Domain EnrichmentARead-onlyIdempotentInspect
Observe a public company domain directly and return bounded website metadata, hashes, injection screening and signed delivery without an expensive enrichment upstream.; 0.003 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, and closed-world, so the safety profile is covered. The description adds useful behavioral context beyond that — fixed price (0.003 USDC), deterministic output, evidence-bounded results, and signed delivery — but omits failure modes, rate limits, or what 'bounded' concretely excludes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Compact and front-loaded with the purpose, which is good. However the stray semicolon and period pattern ('upstream.; 0.003 USDC.') makes the pricing sentence read awkwardly and slightly disrupts structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. Still, with a free-form additionalProperties input schema the description never explains how the target domain is supplied, leaving a real gap for an agent trying to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Parameter count is zero, so per the rubric the baseline is 4. The schema documents no named properties (additionalProperties: true, minProperties: 1), and the description offers no input-shape guidance, but with no parameters to disambiguate there is little room to add value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a concrete verb+resource ('Observe a public company domain directly and return bounded website metadata, hashes, injection screening and signed delivery'), which clearly states what the tool produces. It hints at a sibling distinction ('without an expensive enrichment upstream,' implying phion_enrich), but never names the alternative, so sibling differentiation is only implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the contrast with an 'expensive enrichment upstream,' suggesting this is the lightweight/basic path, but no explicit when-to-use, when-not-to-use, or named alternative (e.g., phion_enrich) is provided. An agent must infer when basic is preferred over the richer siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phion_executeExecute a PHION ServiceAIdempotentInspect
Universal PHION execution gateway: enforce a budget and persistent idempotency, execute one selected PHION service, evaluate acceptance criteria, and return a signed completion envelope. No gateway surcharge; the selected service price applies.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | ||
| dry_run | No | ||
| objective | Yes | ||
| target_service | Yes | ||
| idempotency_key | Yes | ||
| max_budget_atomic | Yes | ||
| acceptance_criteria | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is partly covered. The description nonetheless adds real substance: budget enforcement, persistent idempotency, acceptance-criteria evaluation, a signed completion envelope, and the pricing note ('no gateway surcharge'). It stops short of describing failure/retry behavior or the dry-run semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the gateway purpose and flow, with a pricing clause that earns its place. Dense but not padded; density does cost some readability for a 7-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter, nested-object, no-output-schema mutation tool, the description leaves meaningful gaps: dry-run behavior, input shape, atomic budget units, acceptance-criteria pointer syntax, and what happens when criteria fail are all unaddressed. Both schema and description are thin on parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden, and it only partially compensates. It conceptually covers budget (max_budget_atomic), idempotency (idempotency_key), acceptance criteria, and target_service, but says nothing about the 'atomic' unit format, the JSON-pointer syntax for acceptance criteria, the objective field, or what dry_run does.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('execute one selected PHION service') and frames the tool as a 'Universal PHION execution gateway' that also budgets, idempotency-checks, evaluates acceptance criteria, and returns a signed envelope. That is far more informative than restating the name, but it never explicitly contrasts itself with siblings like phion_preflight or verify, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Execute one selected PHION service' implies this is the terminal step after discovery/selection, but no condition is given for choosing this over phion_preflight or verify, and no prerequisites or exclusions are stated. Usage is deducible from context rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phion_get_serviceGet PHION Service ContractARead-onlyIdempotentInspect
Retrieve one exact canonical PHION service contract by service id or MCP tool name.
| Name | Required | Description | Default |
|---|---|---|---|
| serviceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered. The description adds that the returned contract is 'canonical' and matched exactly, but says nothing about not-found behavior or what the contract contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the verb, scope, and accepted keys; no filler or redundant restatement of the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with rich annotations and no output schema, the description covers the essentials. It could add a word on miss/error behavior, but nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the only property is an undescribed 'serviceId' string, so the description carries the burden. It adds real meaning by stating the key accepts either a service id or an MCP tool name, though it omits id format or case-sensitivity details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Retrieve') plus resource ('PHION service contract') with a scope qualifier ('one exact canonical') that implicitly contrasts with the plural search sibling. It does not name an alternative tool explicitly, so it falls just short of the top band.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states that lookup is by service id or MCP tool name, but gives no guidance on when to use this versus phion_search_services, phion_resolve, or phion_discover. No prerequisities or exclusions are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phion_inferencePHION InferenceARead-onlyIdempotentInspect
Complete a bounded inference microtask and return a directly usable answer with signed delivery.; 0.001 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior. The description adds useful context beyond annotations: determinism, evidence-bounding, signed delivery, and a 0.001 USDC cost. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the main purpose. The semicolon splice and trailing cost note are slightly awkward, but every phrase contributes meaningful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and annotations cover the safety profile, but the description never defines what a 'bounded inference microtask' is or what input shape is expected. Since the schema is an open object, this leaves the input contract underspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero named parameters, so the baseline is 4. The open schema (minProperties 1, additionalProperties true) is not explained, and the description does not specify what fields a microtask should contain, but with no defined parameters the burden on the description is limited.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('complete a bounded inference microtask', 'return a directly usable answer') and adds distinctive traits like 'deterministic, evidence-bounded and signed'. It does not explicitly contrast with sibling phion_* tools, so it stops short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as phion_search, phion_enrich, or phion_execute. The phrase 'bounded inference microtask' implies a use case, but no explicit conditions, exclusions, or selection criteria are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phion_intelligencePHION IntelligenceARead-onlyIdempotentInspect
Search, synthesize and cite evidence in one call, minimizing residual agent work.; 0.012 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive behavior. The description adds useful traits beyond those: 'Deterministic, evidence-bounded and signed' and the fixed cost '0.012 USDC.' This gives an agent expectations about repeatability and output guarantees without contradicting annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the core action, then cost and behavioral guarantees. No filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return format is covered. However, the open input schema and lack of usage guidance mean an agent must infer what to pass and when this tool beats phion_search or phion_inference. Cost and determinism help, but input semantics and selection criteria are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema defines no named parameters, so there is no parameter documentation to augment. The description's 'Search, synthesize and cite evidence' implies a query/topic input, but the open schema (additionalProperties: true) leaves exact keys unspecified; with zero declared parameters, the baseline is 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific composite action: 'Search, synthesize and cite evidence in one call.' This clearly identifies the tool's function and differentiates it from single-purpose siblings like phion_search or phion_inference by emphasizing the combined workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'in one call, minimizing residual agent work' implies use when an agent needs the full search-synthesize-cite pipeline, but it never explicitly names alternatives or states when not to use it. Siblings such as phion_search and phion_inference exist but are not referenced.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phion_preflightCheck PHION Service PreflightBRead-onlyIdempotentInspect
Free service-specific schema, network, budget and economic preflight. It does not execute or charge.
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | ||
| network | No | ||
| intentId | No | ||
| protocol | No | ||
| serviceId | Yes | ||
| budgetUsdc | No | ||
| correlationId | No | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, covering the safety profile. The description adds 'free' and 'does not charge', which is a useful cost behavior not in annotations, but it does not describe return format, permissions, or what a preflight actually validates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler; the non-execution/non-charge statement is front-loaded in second position to clarify scope. The first sentence reads as a compact keyword list but remains efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter tool with a nested object and no output schema, the description is thin. It does not explain what the preflight returns, how to supply the service input, or the meaning of most parameters, leaving substantial gaps for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% with 8 parameters (including a nested 'input' object). The description gestures at 'network' and 'budget' but never maps or explains any parameter, leaving serviceId, intentId, protocol, correlationId, idempotencyKey, and the nested input entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific function: a free service-specific preflight covering schema, network, budget and economic checks. It clarifies the scope and the non-executing nature, which helps distinguish it from phion_execute, though it does not explicitly name sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated — 'It does not execute or charge' hints that this is a safe step before execution, but there is no explicit 'use before phion_execute' or 'when not to use' guidance. The agent must infer the context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phion_resolveResolve an Agent CapabilityBRead-onlyIdempotentInspect
Free provider-neutral capability resolution. Applies eligibility constraints, preserves UNKNOWN and abstains when evidence is insufficient. Advisory only: never authorizes payment or execution.
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | ||
| capability | No | ||
| constraints | No | ||
| executionMode | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the bar is lowered, and the description still adds real behavior: it 'preserves UNKNOWN and abstains when evidence is insufficient' and never authorizes payment or execution. What it does not disclose is what an abstention or UNKNOWN actually returns, which matters for an openWorld resolver.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, no filler, and the core purpose is front-loaded before the behavioral caveats. The dense jargon ('provider-neutral capability resolution') costs a little readability but nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a resolver with a nested constraint object, mutually-exclusive input shapes and a 3-value mode enum, the description covers the conceptual contract but not the invocation surface. With no output schema present, the description should also indicate what a resolution result looks like, and it does not.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters, including an 11-field nested constraints object and an executionMode enum with SHADOW/ADVISORY/PREFLIGHT values. The description says nothing about task vs capability, the anyOf requirement, what constraints are honored, or what the three modes mean, so it does not compensate for the coverage gap at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('capability resolution') and frames scope with 'provider-neutral' and 'free'. It implicitly separates itself from phion_execute by stating it is 'advisory only: never authorizes payment or execution', but it never names or contrasts with the closer siblings phion_discover, phion_search_services or phion_preflight, leaving the boundary fuzzy.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the 'advisory only' clause signals this is the non-committing step and that phion_execute is the committing one. There is no explicit when-to-use statement, no prerequisites (e.g. when to prefer preflight over resolve), and no exclusion rules, so the agent must infer the decision boundary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phion_searchPHION SearchCRead-onlyIdempotentInspect
Search the live web and return normalized source URLs and snippets with signed delivery.; 0.006 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description says 'Search the live web,' implying open-world network access, while annotations declare openWorldHint=false. It also claims 'Deterministic' for a live web search, which is tense. This is an annotation contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in a concise phrase, and the cost plus determinism/evidence-bounded qualifiers add useful constraints. The semicolon fragment and repeated 'signed' wording are slightly awkward, preventing a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema and annotations cover return shape and safety, but the input contract is unspecified and the relationship to sibling search/evidence tools is absent. An agent cannot construct a valid call (it must supply at least one property) nor decide when this tool is appropriate, and the open-world contradiction further undermines trust.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although context reports 0 declared parameters, the schema requires minProperties:1 and allows arbitrary additional properties, so the agent still needs to know what to place in the input object. The description gives no key names, query format, or search options, and the baseline for 0 declared params does not apply because the schema actually demands at least one property.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and resource ('Search the live web') and states the concrete output ('normalized source URLs and snippets with signed delivery'). This clearly distinguishes it from sibling enrichment, inference, and execute tools, which are not search operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to prefer this tool over closely related siblings such as source_backed_search, verified_web_extract, or fetch_evidence, and no exclusions or prerequisites. The only implicit context is 'live web,' which is insufficient for an agent to route correctly among similarly named search/evidence tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phion_search_servicesSearch PHION ServicesARead-onlyIdempotentInspect
Search the complete canonical PHION inventory by task, capability, protocol or maximum advertised price without selecting or charging.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | ||
| maxPrice | No | ||
| protocol | No | ||
| capability | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description usefully reasserts that nothing is selected or charged, but says nothing about result ordering, pagination, or how the canonical inventory changes between calls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, tightly front-loaded with the verb and scope, with zero filler. Every clause (canonical inventory, search facets, no-selection caveat) earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and 0% schema description coverage, the description should do more: it never conveys the shape of results (list of services with prices and capabilities?) or how 'limit' caps them. Adequate for routing, incomplete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the parameter burden. It maps its facets onto query, capability, protocol, and maxPrice ('maximum advertised price'), which is real added value, but leaves the 'limit' parameter unexplained and gives no syntax or combination semantics for the free-text query.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Search) and resource (complete canonical PHION inventory), plus the facets that can be searched: task, capability, protocol, price. The phrase 'without selecting or charging' implicitly separates it from action-oriented siblings like phion_execute or phion_preflight, though no sibling is named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Without selecting or charging' implies this is a pre-selection, discovery-only step, which hints at when to use it. However, it never names the alternative (e.g. phion_get_service for a known service, phion_execute to act) or states explicit exclusions, leaving the agent to infer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
portable_observability_auditBInspect
Validate event identity, timestamps, hash chain and W3C-compatible lowercase nonzero trace/span identifiers; 0.004 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| events | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose meaningful behavioral detail: the specific invariants checked (hash chain, lowercase nonzero W3C trace/span IDs) and the cost (0.004 USDC). However, it says nothing about what a failure yields, whether the operation is read-only, or what the verdict/report looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence with semicolon-separated clauses, front-loaded on the action and the validations performed. The price is appended at the end where it doesn't interrupt the primary meaning; nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a validation tool with no annotations and no output schema, the description covers the input-side contract well but is silent on the return value (pass/fail, error list, per-event verdicts) and on any failure behavior. An agent can invoke it correctly but cannot anticipate what it gets back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With a single parameter at 0% schema description coverage, the schema documents only structure (required event_id, observed_at, trace_id, span_id), not meaning. The description compensates by spelling out what those fields must satisfy — identity validity, timestamp validity, W3C lowercase nonzero trace/span format — which is exactly the semantics the schema omits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Validate') and enumerates the exact things validated: event identity, timestamps, hash chain, and W3C-compatible trace/span identifiers. It is clearly distinguishable from siblings like verify or attest in substance, though it never names a sibling to route against. The trailing price is unusual but doesn't obscure the purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this tool versus alternatives such as verify, attest, or transaction_assurance, and no prerequisites or preconditions are given. The agent must infer usage entirely from the name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preflightCInspect
Free URL safety and reachability check.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden, and it discloses almost nothing beyond 'free' — no auth requirements, no explanation of what makes a URL unsafe, no behavior on failure, and no rate limits. Calling something a safety check without saying what it inspects or returns leaves real ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with no filler; the core action leads. It is efficient, though the terseness edges toward under-specification rather than tightness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with no output schema, the description still omits what a 'safety' verdict consists of and what the caller gets back, which is the main thing an agent needs to decide whether to gate a downstream action on the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single required url parameter, and the description only says a URL is checked. It does not clarify expected form (absolute? scheme? relative paths rejected?), which matters because the schema only declares format: uri with no prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a concrete verb-plus-resource framing: it checks a URL for safety and reachability. That is enough for an agent to know what it does, but it never distinguishes itself from the sibling payment_preflight, which sounds nearly identical in kind.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to call this versus payment_preflight, verify, or the other siblings, and no prerequisites (e.g., authenticate first, call before payment). The agent is left to infer the workflow position of a 'preflight' entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
price_discovery_enginePrice Discovery EngineARead-onlyIdempotentInspect
Derive price bands only from settled observations, never traffic or unsigned quotes.; 0.005 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, idempotentHint=true. The description adds valuable context: the source of evidence (settled observations), the exclusion of non-settled data, determinism, and evidence-bounded nature. It does not contradict annotations and enhances transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise, three short clauses, no fluff. It front-loads the core action and constraints. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters and a simple output (price bands), the description covers the essential input constraints. The '0.005 USDC' is cryptic and might need clarification, and the output schema exists but is not shown, so the description needn't detail return values. Minor ambiguity costs a point.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters and the schema is fully open (additionalProperties: true), so the description doesn't need to explain parameters. The mention of '0.005 USDC' might hint at pricing or stake but is ambiguous, which is a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('derive') and resource ('price bands') and adds constraints ('settled observations, never traffic or unsigned quotes'). However, it doesn't differentiate from siblings like quote_freshness_guard or market_data_snapshot, which could be related.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: when you need price bands from settled data. But it doesn't explicitly state when not to use it or mention alternatives. With 100+ siblings, more explicit routing would help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
proof_of_serviceProof of ServiceCRead-onlyIdempotentInspect
Commit request, authority, execution, payment and delivery evidence in one proof.; 0.005 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description states deterministic, evidence-bounded, and signed, which adds behavioral context, but it contradicts the annotations: it describes a committing action that likely creates a new proof, yet readOnlyHint is true. This is a clear contradiction that undermines trust.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and front-loaded with the core purpose, but it sacrifices clarity for brevity. Every word is purposeful, yet the vagueness reduces its effectiveness. Still, it's a single sentence that is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema, the return value may be covered, but the description provides no details on the input format, the types of evidence accepted, or the exact semantics of 'commit'. The contradiction with annotations and lack of usage guidelines make it incomplete for reliable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description doesn't need to explain any. The schema is open to any properties, but with no fixed parameters, the baseline for parameter semantics is high.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the verb 'Commit' and names a specific resource ('request, authority, execution, payment and delivery evidence'), but it is vague about what the tool actually does—it sounds like a wrapper that bundles several evidence types into one proof, yet the exact operation and output are unclear. It does not distinguish itself from siblings like 'delivery_evidence' or 'fetch_evidence'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no explicit guidance on when to use this tool versus alternatives. It mentions a cost and determinism, but not the conditions under which this should be chosen over similar evidence tools. Agents are left to infer use cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
provider_quality_predictorProvider Quality PredictorBRead-onlyIdempotentInspect
Estimate provider success, quality, latency and cost using transparent sample-aware statistics.; 0.004 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior. The description adds meaningful extra context beyond annotations: it costs 0.004 USDC, is deterministic, evidence-bounded, signed, and uses transparent sample-aware statistics. This gives the agent important expectations about side effects and output characteristics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core action and scope. The phrasing is slightly awkward due to the semicolon and standalone cost token, but every clause adds useful information: purpose, method, fee, and behavioral guarantees.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and annotations cover safety, so the return shape and read-only nature are handled. However, the description does not state what inputs are needed and gives no usage context to distinguish it from sibling estimator/predictor tools. This leaves the agent short of what it needs to invoke the tool correctly from the open input schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema defines no named parameters but permits an open object with additionalProperties and minProperties 1. Although zero defined parameters normally lowers the burden, the open schema means the description should clarify what inputs are expected (e.g., provider identifier or sample data), and it does not. The description names output domains rather than input requirements.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Estimate') and names the resource and scope: provider success, quality, latency, and cost. It is clear about what the tool computes, though it does not explicitly differentiate itself from sibling predictors like sla_risk_predictor or execution_cost_estimator.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to prefer this tool over alternatives, nor does it mention exclusions or complementary tools. A caller is left to infer that this is for provider estimation, but there is no context about selecting it among the many prediction/estimation siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
purpose_bound_consentCInspect
Fail-closed consent bound to ID, subject, recipient, purpose, scope, issue/expiry and checked revocation state; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| consent | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it only partially does so. It discloses that behavior is fail-closed, that revocation state is checked, and that there is a 0.003 USDC cost, but it omits critical behavior such as what happens on success or failure, whether the cost is per call, permissions, side effects, and return semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The text is a single dense fragment with no wasted words, and the price and fail-closed qualifier are front-loaded. However, the lack of a main clause or clear operation makes the structure weak and less useful than a concise but grammatically complete statement would be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given nested required objects, 0% schema description coverage, no annotations, and no output schema, the description is far too sparse to call the tool correctly. It hints at fail-closed consent checking and cost, but does not state the operation, return behavior, or action-parameter semantics, leaving substantial gaps for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for two nested object parameters. It lists semantic fields for the consent object (ID, subject, recipient, purpose, scope, issue/expiry, revocation state) but says nothing about the required action object, which also has required subject, recipient, purpose, and scope fields. The description therefore leaves a major parameter entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a noun phrase describing the consent object and its bound fields, not a verb-driven statement of what the tool actually does. It largely restates the concept behind the name 'purpose_bound_consent' without clarifying whether the tool validates, issues, enforces, or inspects consent. An agent cannot reliably infer the operation from this text alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is provided. The description does not name alternatives or explain under what conditions an agent should choose this tool over sibling guard or preflight tools. It only notes that the tool is fail-closed and costs 0.003 USDC.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quote_freshness_guardQuote Freshness GuardDRead-onlyIdempotentInspect
Quote Freshness Guard; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds only the word 'deterministic' and 'signed', which hints at output behavior but does not explain what 'signed' means or what the assessment entails. No contradictions with annotations, but minimal added value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is only one sentence, but it is unhelpfully vague and lacks front-loaded, actionable information. It does not earn its place because it merely restates the tool's name and adds an opaque phrase.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema, the description fails to explain the input, the meaning of the assessment, or the expected result. For a tool with a single required parameter and a clear purpose implied by its name, this is severely under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one required parameter 'quote' but no description for it. The tool description does not elaborate on what constitutes a 'quote' or how it should be formatted. With zero schema descriptions and no tool-level explanation, an agent cannot correctly construct the input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'deterministic signed assessment over supplied input' is vague and doesn't clearly state what the tool does. It echoes the title without clarifying that it assesses the freshness of a quote. It doesn't distinguish it from sibling tools like data_freshness_certificate or live_data_freshness, leaving an agent unable to determine its unique function.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus any alternative. No context is provided about typical scenarios, prerequisites, or exclusions. The agent has no basis to decide between this and the many freshness-related sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
redirect_callback_validatorBInspect
Validate HTTPS callbacks and redirect host allowlists; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and mostly fails it. It does not say whether validation is a local policy check or performs live network fetches, whether the result is a boolean vs. a report, what failure conditions exist, or any auth/rate-limit context. The cost disclosure (0.002 USDC) is the one genuine behavioral trait it adds.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the validated resource first and the price trailing. No filler, no redundancy, nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained. However, with no annotations and no usage guidance, an agent still lacks the context needed to know when this validator applies and how its result should be interpreted, making the definition only minimally viable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is reported at 100%, so the baseline is 3 even though the description adds no explicit parameter meaning. The phrase 'host allowlists' loosely corresponds to allowed_hosts and 'HTTPS callbacks' constrains requested_url, but the description never clarifies formats, matching rules, or required scheme.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Validate') and two concrete resources ('HTTPS callbacks' and 'redirect host allowlists'), so an agent can tell what it operates on. However, it offers no differentiation from closely related siblings such as webhook_verifier or data_egress_preflight, leaving overlap ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, or alternative is named. The only usage-adjacent information is the price ('0.002 USDC'), which tells the agent nothing about which situations should select this tool over the validators siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reputation_update_receiptReputation Update ReceiptARead-onlyIdempotentInspect
Recommend a bounded auditable reputation update from verified outcome evidence.; 0.004 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds several details beyond the annotations: the operation only recommends rather than applies an update, it is deterministic and evidence-bounded, it produces a signed result, and it costs 0.004 USDC. These are genuinely informative and consistent with the readOnlyHint, idempotentHint, and destructiveHint annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core purpose, followed by cost and behavioral guarantees. No sentence is wasted, and all included details earn their place, even though the semicolon before '0.004 USDC' is stylistically awkward.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers high-level behavior, cost, determinism, and signing, and an output schema exists so return format does not need to be explained. However, it lacks input structure details, usage context, and sibling differentiation, leaving an agent under-equipped to confidently construct a valid call among the large sibling set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema declares no named properties and allows arbitrary properties, while the description's only semantic pointer is 'verified outcome evidence.' The description does not specify what evidence fields, identifiers, or reputation values are expected, so an agent still lacks concrete input-construction guidance despite the nominal zero-parameter count.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Recommend'), a specific resource ('a bounded auditable reputation update'), and a source ('verified outcome evidence'). It reads as a distinct advisory tool rather than an execution or evidence-gathering tool, though it does not explicitly contrast with sibling tools like agent_reputation_evidence or erc8004_reputation_intelligence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'from verified outcome evidence' implies the intended context: use when verified evidence exists and a reputation adjustment is being proposed. However, the description does not state when not to use it, what prerequisites must hold, or how it differs from the many reputation and evidence-related sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resource_cycle_guardResource Cycle GuardBRead-onlyIdempotentInspect
Require per-step identity, token and cost counters; detect repeated nodes/edges and enforce all three hard caps; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| trace | Yes | ||
| policy | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, covering the safety profile. The description adds genuinely new context: the 0.002 USDC cost and the fact that all three caps are hard-enforced. However, it does not say what happens on violation (raise, block, return verdict) or what the response contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the core detection/enforcement behavior front-loaded and the price disclosed last. No filler, though the semicolon-chained phrasing is dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with nested object inputs, two required params and no output schema, the description omits the return contract and the consequence of tripping a cap, which an agent needs to sequence calls correctly. The cost disclosure partially compensates, but the completeness bar is not fully met.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the burden, and it does reasonably: 'per-step identity, token and cost counters' maps to the trace item fields (id, tokens, cost_atomic) and 'all three hard caps' maps to the policy's max_steps/max_tokens/max_cost_atomic. It stops short of naming field names or clarifying the 'atomic' unit semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names concrete actions (require per-step identity/token/cost counters, detect repeated nodes/edges, enforce three hard caps), which tells an agent exactly what this guard does. The cycle-detection and budget-cap combination distinguishes it functionally from neighbors like concurrency_guard, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to invoke this tool, in what pipeline stage, or which alternatives (concurrency_guard, agent_budget_guard, subscription_spend_guard) apply instead. The agent must infer usage entirely from the behavior sentence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retention_deletion_receiptDInspect
Separate requested, operator-completed and evidence-verified retention/deletion states using content and evidence hashes; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| action | Yes | ||
| record | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: no permissions required, no mutation semantics, no idempotency, no cost mechanics beyond mentioning 0.003 USDC, and no explanation of what the three state distinctions mean operationally.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is short but not front-loaded with a purpose—it opens with an abstract list of states before the reader knows what the tool does. The trailing '0.003 USDC' is the only concrete detail and feels disconnected from the rest.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two required nested-object parameters, 0% schema coverage, no annotations, and no output schema, the description is far too thin. An agent cannot determine what to pass or what it will get back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and both parameters are required nested objects (action.kind, record.record_id, record.content_sha256). The description mentions content and evidence hashes but does not map them to parameters or explain valid action kinds, leaving the schema entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description lists concepts (requested, operator-completed, evidence-verified states) but never states a clear verb+resource. 'retention_deletion_receipt' plus a sentence fragment describing states makes it hard to tell whether this creates a receipt, verifies one, or reports status. It does not differentiate itself from siblings like delivery_evidence or verify.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no when-not-to-use, and no mention of alternatives such as delivery_evidence, attest, or verify. The agent must infer the situation from the description alone, which is not possible with confidence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rwa_asset_due_diligenceBRead-onlyIdempotentInspect
Fail-closed issuer, contract, jurisdiction, custody and documentation coverage with independently fetched evidence; 0.019 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| asset | Yes | ||
| source_urls | No | ||
| declared_facts | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotent, read-only, open-world, non-destructive behavior, so the bar is lower. The description still adds two valuable facts annotations cannot convey: it is fail-closed (rejects rather than degrades when evidence is missing) and it carries a cost of 0.019 USDC, which matters for a paid/x402 call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with zero filler, and the fail-closed posture is front-loaded. It is telegraphic to the point of being jargon-dense, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists and the input schema is essentially opaque (a bare nested object), so the description would need to carry more weight. It covers behavior and price but omits what evidence is returned, what a failed/closed result looks like, and what the undocumented parameters must supply.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and 'asset' is a nested object with no documented fields, yet the description says nothing about what asset, source_urls, or declared_facts should contain. 'Independently fetched evidence' hints at the tool's posture, not at parameter shape, so the three undocumented parameters are never compensated for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (RWA asset due diligence) and enumerates the exact coverage dimensions: issuer, contract, jurisdiction, custody, documentation. That is far more concrete than a tautology, but it never distinguishes itself from close siblings such as rwa_compliance, rwa_sanctions_screening_evidence, or rwa_transaction_assurance, so an agent cannot route between them from the text alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, and no alternative sibling is named. 'Fail-closed' implies it is meant for gating decisions, but the agent is left to infer the trigger condition and the relationship to the other rwa_* tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rwa_complianceCRead-onlyIdempotentInspect
Fail-closed RWA transfer decision requiring explicit jurisdiction, KYC, allowlist, limit and transfer-window checks; 0.012 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| asset | Yes | ||
| policy | Yes | ||
| source_urls | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotency, non-destructiveness and open-world access, so the safety profile is not the description's burden. The description does add two genuinely non-structured traits: fail-closed denial semantics and a 0.012 USDC charge, which matter for an agent deciding whether to invoke a paid gate. It still omits what a denial looks like and whether the call is blocking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the most important trait ('fail-closed') front-loaded and no filler. It is efficient, though the semicolon-separated list of checks and trailing price make it read as a packed clause chain rather than clearly structured guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A paid, open-world, two-nested-object tool with no output schema and zero schema descriptions is significantly under-documented: the description never explains the shape of 'asset' or 'policy', what source_urls is for, or what the decision result contains. For the complexity indicated, this is not enough to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and both required parameters ('asset', 'policy') are untyped nested objects, so the schema provides no semantics at all. The description lists the check categories (jurisdiction, KYC, allowlist, limit, transfer-window), which weakly hints at what 'policy' must contain, but never explains the asset object or the role of source_urls.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the domain (RWA transfer compliance) and enumerates the check categories it covers, but the operative verb is only implied ('decision') and it gives no signal distinguishing it from siblings like rwa_transaction_assurance or transaction_assurance. An agent can guess its role but cannot confidently route between the overlapping RWA/assurance tools from this text alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use statement, no prerequisites, and no named alternatives. 'Fail-closed' hints that it is a gate, but nothing tells the agent at which point in a transfer flow to call it versus rwa_sanctions_screening_evidence or payment_preflight.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rwa_corporate_actionsCInspect
Detect and sign material RWA changes in NAV, income, maturity, redemption or freeze state; 0.012 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| asset | Yes | ||
| source_urls | No | ||
| current_state | Yes | ||
| previous_state | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full behavioral burden. It discloses the 0.012 USDC fee, which is useful, but 'sign' implies an on-chain/write action and key requirements that are never explained, and nothing is said about reversibility or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the core action front-loaded and the cost appended. It is efficient and waste-free, though arguably too terse for a four-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with nested object parameters, 0% schema coverage, no output schema, and no annotations, the description leaves too much unexplained: parameter shapes, signing requirements, and return behavior are all absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters (3 required), including nested objects like asset/previous_state/current_state. The description only hints at attributes ('NAV, income, maturity, redemption or freeze state') that likely live inside those state objects, leaving source_urls and the nested shapes undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific verbs (detect, sign) and a concrete resource domain (material RWA changes in NAV, income, maturity, redemption, freeze state), which is far more informative than the bare name. It does not, however, distinguish itself from siblings like rwa_nav_reserve or rwa_compliance, so the agent must infer boundaries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no alternatives are named despite dense sibling overlap (rwa_nav_reserve, rwa_compliance, rwa_asset_due_diligence). The only routing-adjacent info is the price tag, which is not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rwa_sanctions_screening_evidenceBInspect
Evidence-based literal subject screening against observed official US/EU sanctions sources; 0.015 USDC. Not legal clearance or wallet attribution.
| Name | Required | Description | Default |
|---|---|---|---|
| subject | Yes | ||
| source_urls | No | ||
| jurisdictions | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the cost (0.015 USDC), the evidence nature (observed official sources, literal screening), and a limitation (not legal clearance). However, it omits return format, permissions, whether it is read-only, and what evidence is produced, leaving moderate gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence with semicolon-separated clauses that are front-loaded with the core purpose, cost, and limitation. Every clause earns its place and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a nested input schema, no output schema, and no annotations, the description is too thin. It lacks any explanation of return values, error behavior, or how the subject and source_urls parameters are structured, leaving an agent without enough context to invoke it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% with three parameters, including a nested subject object and a source_urls array. The description provides no additional meaning for any parameter—no guidance on subject fields, jurisdiction usage, or source_urls purpose—so it fails to compensate for the undocumented schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'literal subject screening against observed official US/EU sanctions sources.' It is clear what the tool does and distinguishes itself from the broader compliance siblings by emphasizing evidence-based, literal screening. However, it does not explicitly name sibling alternatives to further sharpen differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Not legal clearance or wallet attribution' sets a boundary by stating what the tool is not, which implies it should be used for evidence-based screening rather than legal opinions or wallet attribution. But it gives no explicit when-to-use guidance relative to siblings like rwa_compliance or verify, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rwa_transaction_assuranceCInspect
Verify an RWA asset, payment, delivery and transfer restrictions; 0.020 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| asset | Yes | ||
| policy | Yes | ||
| source_urls | No | ||
| transaction | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, yet it only discloses the price (0.020 USDC). It does not say whether this is a read-only or state-changing operation, what determines pass/fail, whether verification is deterministic, or how failures are surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, and the cost is placed at the end where it is easy to scan. It is efficient, though the brevity reflects under-specification as much as discipline.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three required nested-object parameters, no annotations, no output schema, and zero schema coverage, the description leaves critical context missing: no input shape guidance, no indication of what a successful verification returns, and no routing versus sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With four parameters, three required, 0% schema description coverage, and nested objects, the description must compensate and does not. 'Asset, payment, delivery and transfer restrictions' is a loose gloss that never explains what belongs in the asset, transaction, or policy objects, nor the role or cap of source_urls.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Verify') and enumerates the resources covered: RWA asset, payment, delivery, and transfer restrictions. This is more than a restatement of the name, though it never distinguishes itself from close siblings like rwa_compliance or transaction_assurance, which appear to operate in the same space.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus rwa_compliance, rwa_asset_due_diligence, transaction_assurance, or the preflight family. The agent is left to infer selection from the name alone, and the description gives no conditions, prerequisites, or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
schema_normalizeBRead-onlyIdempotentInspect
Safe alias normalization that refuses ambiguous canonical/alias collisions and never changes payment values; 0.001 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | ||
| target_service | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds real behavioral context beyond annotations: it refuses collisions (a failure mode), is scoped to aliases, and explicitly preserves payment values (a data-safety guarantee). Annotations already cover readOnly/idempotent/non-destructive, so this layered disclosure is genuinely useful. Rate limits and error format remain unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense, front-loaded clause covering scope, refusal behavior, payment-safety guarantee, and cost. No filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two required params at 0% schema coverage, a nested payload object, and no output schema, the description omits the input contract entirely. The correctness-relevant facts about how target_service and payload are consumed are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the two required params (target_service, payload). The description says nothing about what goes into target_service or payload shape, or how target_service selects the schema to normalize against. With nested objects and a 0% coverage, the description should compensate but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific operation ('safe alias normalization') plus distinctive behavioral traits (refuses ambiguous canonical/alias collisions, never changes payment values). It's clearly distinguished from the paid/free sibling schema_normalize_free and other verification tools, though the phrase 'safe alias normalization' is domain-specific enough that its exact effect on payloads is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this vs schema_normalize_free or any other sibling. The '0.001 USDC' cost is disclosed, which implies a paid variant exists, but no selection criteria or exclusions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
schema_normalize_freeBInspect
Free, rate-limited payload validation and safe key normalization before payment. Payment-critical values are never changed.
| Name | Required | Description | Default |
|---|---|---|---|
| payload | Yes | ||
| target_service | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden; it does disclose genuine behavior beyond the schema: free tier, rate limiting, non-mutation of payment-critical values. But it omits error/invalid-payload behavior, what the rate limit actually is, and whether normalization mutates or returns a new payload.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the key guarantee front-loaded after the purpose clause; no filler. Slightly dense with unattributed modifiers ('Free, rate-limited') but appropriately sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations, with a nested free-form payload object and an undocumented required target_service string. An agent cannot reliably construct a valid invocation from this description alone; payload shape, service identifiers, and failure behavior are all missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%: the nested 'payload' object has no field documentation, and 'target_service' is never mentioned in the description at all, so its expected values are entirely unknown. The description adds some intent ('payment-critical values are never changed') but does not compensate for the uncovered, required parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action pair (validation and safe key normalization) on a specific resource (payload) with scope (before payment). The adjectival 'Free, rate-limited' implicitly distinguishes it from the sibling schema_normalize, but the distinction is left for the agent to infer rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear timing context ('before payment'), which implies when to reach for it, but never names the alternative schema_normalize or states when to use the paid variant instead. No exclusions or prerequisites are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
secret_redaction_preflightCInspect
Detect and redact common secret indicators before transmission; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It mentions the action and a cost ('0.002 USDC'), but omits critical traits: whether the payload is modified in place, what counts as a 'secret indicator', whether redaction is reversible, required permissions, and any side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that states the core action, the timing constraint, and the cost. Every element earns its place, and there is no verbosity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, an output schema, and many similar preflight siblings, this one-line description is too sparse. It lacks usage direction, behavioral details about redaction or mutation, and any input expectations beyond the name, leaving the agent underinformed for selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is reported as 100% and the parameter count is 0, so the schema is treated as fully documenting available parameters. The description adds no parameter-level meaning, which is acceptable given the baseline, but it also does not clarify the required 'payload' input at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb pair ('Detect and redact') and resource ('common secret indicators') along with context ('before transmission'). It is clear what the tool does, but it does not differentiate itself from related preflight siblings such as data_egress_preflight or payment_preflight.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'before transmission' implies a usage context, but there is no explicit guidance on when to choose this tool over alternatives, no exclusions, and no prerequisites. With many preflight siblings, this leaves selection ambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
service_failover_selectorService Failover SelectorCRead-onlyIdempotentInspect
Service Failover Selector; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds 'deterministic' and 'signed,' which are mildly useful behavioral traits not present in the annotations, but it does not clarify what 'signed' means or what the assessment returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded, but the brevity is under-specification rather than efficiency. It spends one clause restating the title and a second clause on a phrase so generic that it conveys little operational meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complex tool name, the large sibling context, and a required but undocumented 'candidates' parameter, the description is far from complete. An output schema exists, but the description still needs to explain the assessment basis, the shape of candidates, and when this selector is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema requires 'candidates' but provides no property definitions, types, or constraints. The description never mentions 'candidates' or explains what shape the input should take, so an agent has no way to construct a valid call beyond guessing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description restates the tool name ('Service Failover Selector') and adds the vague phrase 'deterministic signed assessment over supplied input.' It does not explain what the selection is based on, what 'service failover' means in practice, or how this differs from sibling tools like payment_route_selector or try_service.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives. The sibling list contains many selection and assessment tools, but the description never names a condition, prerequisite, or exclusion that would help an agent route to this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
service_gap_detectorService Gap DetectorBRead-onlyIdempotentInspect
Detect capability shortages using verified demand and verified provider supply.; 0.004 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, and non-destructive behavior. The description adds genuinely useful traits beyond that: deterministic, evidence-bounded, and signed, plus a fixed cost. It does not explain what 'signed' implies for the output, but the annotation coverage is strong enough that this is not a major gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the purpose, but the structure is fragmented: a stray semicolon and an abrupt price fragment ('; 0.004 USDC.') disrupt readability. No words are wasted, but the formatting is clumsier than it should be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, so return-value documentation is not required, and the annotations cover the safety profile. Still, the open input schema and lack of exact request-body guidance leave an invocation gap, and there is no usage guidance to help the agent choose this tool over related siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero named parameters, so the schema itself provides little to work with. The description adds conceptual meaning by naming the inputs: 'verified demand and verified provider supply.' However, the input schema's minProperties: 1 with additionalProperties: true still leaves some ambiguity about the exact expected object shape.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a concrete action and resource: 'Detect capability shortages' and specifies the method ('using verified demand and verified provider supply'). This makes the core purpose clear, though it does not explicitly distinguish itself from siblings such as capability_verification or market_demand_predictor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, prerequisites, or alternatives are provided. With a large sibling set of demand, supply, verification, and prediction tools, the agent is left to infer when this detector should be selected instead of being told.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
service_sla_attestationCInspect
Sign availability and latency calculations from bounded samples; 0.004 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It mentions a cost ('0.004 USDC'), which is useful, but does not state whether signing is irreversible, what permissions/identity are required, latency/rate behavior, or what 'bounded samples' constrains. Output schema exists, so return format needn't be described, but safety and irreversibility are missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with no waste, and the cost is front-loaded in a parenthetical. It is appropriately sized, though it could be slightly more front-loaded on the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a signing/attestation tool with no annotations and a required-samples+policy contract, the description is too thin. It omits prerequisites, irreversibility, identity requirements, and any routing rationale, leaving meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parameters samples and policy are already documented in the schema and baseline is 3. The description adds no syntax, format, or constraint meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (sign) and resource (availability/latency calculations), distinguishing it from generic siblings like 'attest' or 'verify'. However, the scope is vague: 'bounded samples' and 'sla attestation' aren't fully explained, and an agent can't easily tell why this differs from 'attest' without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance. It does not explain when an agent should choose service_sla_attestation over attest, verify, or try_service, nor what prerequisite state (samples, policy) must exist first.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
signed_result_comparatorBInspect
Compare agent results and sign agreement or conflict; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses a cost (0.003 USDC) and a mutation-like action (signing), but says nothing about permissions, reversibility, idempotency, or what the signature means. For a signing/payment operation with zero annotation coverage, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence front-loads the verb and resource and appends the cost. No waste, no filler, and the essential transaction trait (price) is not buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool involves signing and a fee, and it has an output schema, so return values need not be explained. But the absence of any behavioral context for a paid signing operation leaves an agent under-informed about prerequisites and consequences, making this only minimally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is reported at 100%, though the input schema itself is terse. With high schema coverage, the baseline is 3 even though the description adds no parameter-level detail beyond the implied 'results' input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Compare') and resource ('agent results') and adds the action taken ('sign agreement or conflict'), which distinguishes it from neighbors like conflict_resolution and verify. It is clear what the tool does, though it doesn't explicitly name which sibling it is not to be confused with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use, when-not-to-use, or alternative routing guidance. The description mentions the output artifact (signed agreement/conflict) but not the scenario that calls for this comparator over related tools such as conflict_resolution.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sla_risk_predictorSLA Risk PredictorARead-onlyIdempotentInspect
Estimate SLA failure probability from bounded success and latency evidence.; 0.004 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey read-only, idempotent, non-destructive, and closed-world behavior. The description adds valuable traits beyond annotations: determinism, evidence-boundedness, signed output, and a cost of 0.004 USDC. None of these contradict the annotation hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core purpose. The '0.004 USDC' fragment is slightly awkwardly attached with a semicolon, but overall every clause earns its place without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists and annotations cover safety, so the description need not explain return values or mutability. However, the description does not clarify what 'bounded' means, how to provide evidence, or what 'signed' implies operationally, leaving some practical ambiguity for an agent constructing a call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero formally defined parameters and 100% schema coverage, the baseline is 4. The description still adds semantic guidance by indicating that inputs should represent bounded success and latency evidence, which is meaningful for an open-schema object.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Estimate') with a clear object ('SLA failure probability') and names the input evidence ('bounded success and latency evidence'). It distinguishes itself from generic predictors like market_demand_predictor or transaction_risk_score, though it does not explicitly contrast with sibling service_sla_attestation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance, exclusions, or references to alternative tools are provided. The description implies a use case but does not tell an agent when to choose this tool over e.g. service_sla_attestation or provider_quality_predictor.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
source_backed_searchBInspect
Literal search over supplied public sources with citations and hashes; 0.004 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| source_urls | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It discloses the output nature (citations and hashes) and a monetary cost of 0.004 USDC, which is useful behavioral context. However, it does not mention whether the operation is read-only, any rate limits, or error behavior. The cost disclosure is valuable, but the lack of side-effect or safety information leaves some gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, information-dense sentence that includes the core operation, key output features, and cost. No filler words; every element adds value. It is appropriately front-loaded and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's low complexity (two simple parameters, no output schema), the description covers the essentials: what it does, what it returns, and its cost. However, it omits details like the expected output format (e.g., list of results, ordering), any limitations (beyond max 5 sources in schema), and how errors or invalid URLs are handled. For a simple tool this is acceptable, but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description should compensate by explaining what query and source_urls mean. It only implies that source_urls are supplied public sources and that query is a literal search string. It does not specify how to format the query (e.g., single string vs. multiple terms), what constitutes a valid source URL, or any constraints beyond those in the schema. This provides minimal additional meaning over the field names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise action ('search') with a specific resource ('supplied public sources'), and adds distinctive details: 'literal' search, citations, hashes, and a fixed cost. This clearly distinguishes it from sibling tools like verified_web_extract or journey_verify, which focus on extraction or verification rather than user-supplied source search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention any exclusions, prerequisites, or conditions that would lead an agent to choose this over other search-like tools. All contextual selection must be inferred from the tool name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
subscription_spend_guardSubscription Spend GuardCRead-onlyIdempotentInspect
Bind one recurring charge to approval, merchant, due time and per-charge/period budgets; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| usage | Yes | ||
| policy | Yes | ||
| subscription | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered by structured fields and the description need not repeat it. The description adds the per-invocation price (0.003 USDC) and the constraint dimensions it evaluates, which is useful, but it does not say whether 'binding' persists any state, what happens on a policy violation, or what the call returns - notably more important here given the read-only annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tightly packed sentence that front-loads the action and ends with the cost; no filler. It is telegraphic and reads more like a label than prose, but every clause carries information and nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-required-parameter tool with fully nested schemas, 0% schema description coverage, and no output schema, the description is too thin: it never explains the undocumented 'usage' input, the atomic amount units, or the outcome of a call (approve/reject/throw). Annotations cover safety but not the semantics an agent needs to construct a valid invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and all three parameters are nested objects with only 'period_limit_atomic' and 'next_charge_atomic' named, no units, and an entirely opaque 'usage' object. The description partially compensates by naming concepts (merchant, due time, per-charge/period budgets) and the atomic-unit convention is hinted by the price, but it does not map those concepts to the subscription/policy/usage objects or explain how to populate 'usage' at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (one recurring charge/subscription) and the constraint surface it attaches (approval, merchant, due time, per-charge and period budgets), which is more specific than the title alone. However, the verb 'Bind' is unusual and the text never distinguishes this tool from budget/approval siblings such as agent_budget_guard, human_approval_policy, or payment_preflight, so an agent cannot route between them from the description alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of an alternative tool. With ~60 sibling guard/preflight tools, the absence of any routing signal (e.g. 'use this instead of payment_preflight for recurring charges') leaves selection to guesswork.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sybil_reputation_guardSybil Reputation GuardCRead-onlyIdempotentInspect
Combine multiple bounded Sybil signals while abstaining without independent evidence.; 0.004 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds some behavioral context: it is deterministic, evidence-bounded, and signed, and it mentions an abstention behavior. However, it does not explain what 'signed' means in practice (e.g., cryptographic signature, attestation format) or what happens when it abstains. The phrase '0.004 USDC' is unexplained and could be a cost, a reward, or a threshold, which is a transparency gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short, which is good for conciseness, but the structure is choppy: 'Combine multiple bounded Sybil signals while abstaining without independent evidence.; 0.004 USDC. Deterministic, evidence-bounded and signed.' The semicolon and sentence fragments make it read like a list of tags rather than a coherent description. Every phrase is dense, but the lack of connective structure hurts clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the open schema, the presence of an output schema, and the cryptic domain language, the description is not complete enough for an agent to know how to invoke the tool correctly. It does not explain what inputs are expected, what the output contains, what 'abstaining' means for the return value, or what the '0.004 USDC' refers to. The annotations cover safety but not invocation semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema is essentially open (minProperties 1, additionalProperties true) with no named parameters, so the schema provides no semantic guidance. The description's mention of 'multiple bounded Sybil signals' implies the tool accepts a collection of signals, which is the only parameter-level meaning available. With zero named parameters, the baseline is 4, and the description does not actively mislead, though it is thin.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Combine multiple bounded Sybil signals while abstaining without independent evidence' is cryptic and does not clearly state what the tool does or what resource it acts on. It mentions combining signals and abstaining, but the meaning of 'bounded Sybil signals' and 'abstaining' is opaque without domain knowledge. The title 'Sybil Reputation Guard' suggests a reputation/identity protection role, but the description does not clearly distinguish it from siblings like agent_reputation_evidence or erc8004_reputation_intelligence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives. The description implies a decision-making or aggregation behavior ('combine', 'abstain') but does not state the conditions under which an agent should invoke it, nor does it name any alternative tools. The sibling list contains many reputation-related tools, and the description provides no routing information.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
task_cancel_assuranceTask Cancel AssuranceCRead-onlyIdempotentInspect
Task Cancel Assurance; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds the traits 'deterministic' and 'signed' beyond those annotations, but it does not explain what signing means here or what the assessment affects. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief, but it is under-specified rather than meaningfully concise. One sentence of generic wording does not earn its place because it adds no actionable information beyond the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having an output schema and annotations, the description leaves the core semantics unknown: what task cancellation assurance is, what inputs are expected, and how this relates to the many cancellation/transaction siblings. An agent cannot reliably select or invoke this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema exposes only required field names 'task' and 'side_effects' with no property definitions, and the description refers only to 'supplied input.' It does not explain what a task is, what side effects means, or what shape either value should take, leaving required inputs effectively undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description essentially restates the tool name ('Task Cancel Assurance') and adds only the vague phrase 'deterministic signed assessment over supplied input.' It never identifies the specific action, resource, or cancellation context, and with siblings like task_checkpoint_evidence and transaction_assurance, an agent cannot tell what this tool is for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus alternatives such as task_lease_guard, transaction_assurance, or transaction_recovery. No scenarios, prerequisites, or exclusions are mentioned, so the agent must guess based on the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
task_checkpoint_evidenceTask Checkpoint EvidenceCRead-onlyIdempotentInspect
Task Checkpoint Evidence; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, and non-destructive behavior, lowering the burden. The description adds that the result is deterministic and signed, which is useful context beyond the annotations, but it does not clarify side effects, requirements, or output behavior beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with no filler and the key behavioral trait (deterministic signed assessment) is front-loaded. It is concise, though the brevity comes at the cost of operational detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even with annotations and an output schema, the description leaves the core question unanswered: what exactly is being checkpointed and assessed? With an unconstrained input schema and a long list of sibling evidence tools, an agent has too little information to select and call this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema requires 'checkpoint' and allows additional properties, but neither the schema nor the description explains what a checkpoint is, what shape it should take, or what other inputs are accepted. The phrase 'supplied input' does not compensate for the absence of any parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The phrase 'deterministic signed assessment over supplied input' indicates the tool evaluates input and produces a signed result, but it never states a specific action verb or what distinguishes this evidence tool from the many similar sibling evidence tools. It leans heavily on the name itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus an alternative such as task_cancel_assurance, agent_task_handoff_receipt, or other evidence/assessment tools. No context, exclusions, or selection criteria are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
task_lease_guardTask Lease GuardCRead-onlyIdempotentInspect
Task Lease Guard; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly and non-destructive behavior. The description adds that the result is deterministic and signed, which is useful beyond annotations, but it does not explain signing mechanics, required auth, or failure semantics. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short, but the phrase 'Task Lease Guard;' is a near-duplicate of the title and the remaining clause carries little content. This is under-specification rather than intentionally concise front-loading.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists and annotations cover safety, the description omits the core domain semantics (what a lease is, what assessment means), parameter meaning, and selection context. An agent cannot reliably decide when or how to invoke this.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema names required fields 'lease' and 'requester' without types or property definitions. The description only says 'over supplied input,' so it adds no meaning to those fields. The 100% schema-description coverage signal is vacuous because there are no defined properties.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description mostly restates the tool name ('Task Lease Guard') and adds only the generic claim that it performs a 'deterministic signed assessment over supplied input.' It does not say what a task lease is, what property of the input is assessed, or why an agent would call this rather than a sibling such as task_cancel_assurance or resource_cycle_guard.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no exclusion context, and no reference to any sibling tool or alternative. An agent must infer entirely from the name when a 'lease guard' is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tool_call_policy_guardCInspect
Bind a proposed tool call to declared intent and execution policy; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| call | Yes | ||
| intent | Yes | ||
| policy | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It does not disclose whether the tool blocks execution, merely validates, is read-only, or what it returns; the '0.002 USDC' cost hint is the only behavioral detail, and it is terse.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence front-loads the core action and includes the cost. It is efficient, though the compression is part of what leaves the contract underspecified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-required-parameter, nested-object tool with no annotations and 0% schema coverage, the description is far too thin. Even noting that an output schema exists, the input contract and the guard's effect are not conveyed, so the definition is not sufficient to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and all three parameters are untyped nested objects ('call', 'intent', 'policy'). The description names the concepts but adds no structure, expected shape, or required fields for any of them, so the agent must guess the object contracts.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Bind') and objects ('proposed tool call', 'declared intent', 'execution policy'), which is more specific than the name alone. However, 'Bind' is abstract jargon and the tool's actual effect (validation? logging? blocking?) is left ambiguous, so an agent cannot fully distinguish it from siblings like delegation_scope_guard or preflight.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is given. With sibling tools such as preflight, payment_preflight, mcp_manifest_firewall and delegation_scope_guard all plausibly overlapping with policy enforcement, the agent has no signal about which guard to invoke in which scenario.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tool_capability_drift_monitorCInspect
Detect tool capability or manifest drift; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden, and it discloses almost nothing: no indication that this is a read-only comparison, no determinism/idempotency statement, no auth requirements, no note on what a 'drift' finding triggers. The only extra fact is the 0.002 USDC price, which is useful cost context but not behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single front-loaded clause with no filler; the purpose comes first and the price is appended compactly. It is arguably too terse rather than verbose, which is an under-specification problem rather than a conciseness one.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not required, but for a two-input comparison tool the description never clarifies what the inputs represent or what class of drift it covers. Against ~60 siblings with overlapping manifest/tool-integrity functions, the definition leaves the agent without enough to invoke it correctly or confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 100%, the schema is expected to document the two required inputs, so a baseline of 3 is appropriate. The description adds nothing about what 'baseline' and 'current' should contain (manifests? capability snapshots? versions?), which is the one thing a drift tool most needs to clarify.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (detect) and object (tool capability or manifest drift), so the purpose is more than a tautology. However, it gives no differentiation from close siblings such as manifest_version_diff, mcp_manifest_firewall, or tool_output_firewall, which an agent must choose between. The phrase 'capability or manifest drift' is broad enough to overlap with several of them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no named alternative. The agent is left to infer from the name alone that this compares a prior state against a current state, and nothing tells it when manifest_version_diff would be the better pick.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tool_output_firewallCInspect
Inspect untrusted MCP/tool output before it enters agent context or triggers an action; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| output | Yes | ||
| policy | No | ||
| source | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full disclosure burden. It discloses cost (0.002 USDC), which is useful, but says nothing about what happens on unsafe output (block vs. report a verdict), any permission/auth needs, or latency — all critical for a 'firewall' tool whose name implies enforcement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence packs purpose, timing, and price with no wasted words. It is efficient, though the terseness borders on under-specification rather than pure conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values, but with no annotations and zero parameter documentation, the agent cannot construct a valid 'policy' or 'source' object. For a tool with nested-object inputs and a high-stakes enforcement role, the description leaves substantial gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 3 parameters, including nested 'policy' and 'source' objects, so the description must compensate and does not. It only loosely implies the meaning of the required 'output' parameter and gives no indication of what 'policy' or 'source' should contain.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Inspect') and resource ('untrusted MCP/tool output'), giving the agent a clear sense of the operation. It does not explicitly differentiate itself from sibling guard tools like memory_write_guard or delegation_scope_guard, but the resource is specific enough to distinguish it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'before it enters agent context or triggers an action' gives a clear timing condition for invocation. However, it names no alternatives and gives no when-not guidance, leaving the agent to infer when this is preferable to related siblings such as preflight or try_service.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tool_result_schema_validatorTool Result Schema ValidatorCRead-onlyIdempotentInspect
Tool Result Schema Validator; deterministic signed assessment over supplied input.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive hints. The description adds a little beyond them by noting the result is 'deterministic' and 'signed,' which are useful behavioral traits. However, it does not explain what signing entails, what the assessment covers, or any other operational behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short, but the opening phrase 'Tool Result Schema Validator;' is redundant with the name and title. The remaining phrase is compact but vague. There is no wasted length, yet the first half fails to earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite rich annotations and an output schema, the description leaves the core behavior unstated: that a result is checked against a schema and a signed assessment is produced. An agent cannot confidently infer the exact purpose or expected input semantics from the description alone. This is more than minor—it is the central missing piece.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is high, so the schema already carries baseline parameter meaning. The description adds no detail about how 'result' and 'schema' interact, but per the rubric the baseline of 3 applies when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins by repeating the tool title almost verbatim, then adds only 'deterministic signed assessment over supplied input.' It never states the actual operation—that it validates a result against a supplied schema—and offers no differentiation from siblings like signed_result_comparator or schema_normalize. This reads as a tautology with a vague qualifier.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, no exclusions, and no implied context such as 'use when you need to verify a result conforms to a schema.' The agent must infer usage entirely from the tool name and parameter names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transaction_assuranceDRead-onlyIdempotentInspect
End-to-end settlement, timestamp, delivery-hash and deterministic acceptance assurance; 0.015 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| offer | Yes | ||
| intent | Yes | ||
| payment | Yes | ||
| delivery | Yes | ||
| acceptance | No | ||
| transaction_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds 'deterministic acceptance' and a 0.015 USDC cost, but does not explain what is produced (attestation? receipt?), what determinism guarantees, or how the cost is charged. It also doesn't contradict the read-only hint, though 'settlement assurance' sits uncomfortably close to implying mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single compact line with no filler, so it is concise in a literal sense. However, it reads as a marketing tagline with the price appended rather than a front-loaded functional statement, so the brevity comes at the cost of usefulness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity tool with six params (5 required, nested objects), 0% schema coverage, no output schema, and no annotations explaining what the operation yields, the description is inadequate. An agent cannot construct a valid call or predict the result from this text.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are six parameters, five required, with nested object types (intent, offer, payment, delivery) and an acceptance array (maxItems 32). The description names a few concepts (settlement, timestamp, delivery-hash, acceptance) that vaguely map to delivery/acceptance but leaves transaction_id, intent, offer, and payment completely undefined. For 0% coverage, the description must compensate and it does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a dense noun-phrase list ('settlement, timestamp, delivery-hash and deterministic acceptance assurance') plus a price, with no verb stating what the tool actually does to those inputs. It neither confirms it verifies/records/settles a transaction nor distinguishes itself from close siblings such as rwa_transaction_assurance, delivery_evidence, or payment_receipt_reconciler.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not-to-use, prerequisites, or sibling routing is provided. Against ~60 siblings including rwa_transaction_assurance and delivery_evidence, the agent gets zero guidance on when this specific tool is the correct choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transaction_recoveryDRead-onlyIdempotentInspect
Failure classification and non-repeating recovery plan with validated attempt history; 0.010 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| policy | Yes | ||
| payment | Yes | ||
| attempts | No | ||
| assurance | Yes | ||
| transaction_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=false, destructiveHint=false — a strong safety profile. The description adds only 'validated attempt history' and the price, but does not explain what 'failure classification' returns, what the recovery plan contains, whether payment is charged on failure, or how the 0.010 USDC charge is incurred. The pricing is the one genuinely new fact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single short line, so there is no bloat, but the structure is a run-on noun phrase with the price appended after a semicolon. It is concise but not front-loaded with an action; the billing detail sits where the core behavior should be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 4 required parameters (three nested objects), 0% schema descriptions, no output schema, and no explanation of the failure classification or recovery plan semantics, the definition is far too thin. An agent cannot determine valid inputs or interpret the result from this description alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 5 parameters, 4 of which are required, including three nested objects (policy, payment, assurance). The description mentions 'attempt history' (mapping to attempts) but says nothing about transaction_id, assurance, payment, or policy. At 0% coverage the description must compensate and it does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description says 'Failure classification and non-repeating recovery plan with validated attempt history' — a noun phrase, not a verb+resource statement. It never states what action the tool performs (e.g., 'classify a failed transaction and produce a recovery plan'). The '0.010 USDC' price tag is a billing note, not a purpose statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool vs. siblings like interrupted_task_recovery, payment_diagnose, or transaction_assurance. The phrase 'non-repeating recovery plan' hints at avoiding duplicate attempts, but no explicit when/when-not or alternative-tool routing is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transaction_recovery_v2Transaction Recovery v2ARead-onlyIdempotentInspect
Recommend bounded recovery and reconcile ambiguous payments before retry.; 0.005 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive, and the description reinforces compatibility by framing the action as 'Recommend' rather than execute. It adds non-obvious traits: deterministic behavior, evidence-bounded reasoning, signed output, and a 0.005 USDC cost, none of which appear in the schema. No contradiction with readOnlyHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
At roughly 18 words, it is extremely economical and front-loads the main purpose before behavioral traits. The standalone '0.005 USDC' fragment and clipped 'Deterministic, evidence-bounded and signed' read as bullet points, which is slightly awkward but not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple to invoke with no required parameters, has an output schema, and has strong annotations, so the description does not need to explain return values or safety. It supplies the missing selection context—when to use it, for what input condition, and what behavior to expect—but leaves 'bounded', 'evidence-bounded', and 'signed' unexplained for less domain-aware agents.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero named parameters and 100% schema coverage, the schema imposes no burden, so the description does not need to explain parameters. It adds an operational detail (the USDC cost) but no input syntax, which is acceptable because the input schema is open and effectively parameterless.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a concrete action ('Recommend bounded recovery') and adds a second specific behavior ('reconcile ambiguous payments before retry'). It is distinguishable from siblings like transaction_recovery or payment_diagnose because it scopes to ambiguous payments and pre-retry recommendations, though it relies on jargon such as 'bounded' and 'signed' and does not name a comparison sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Before retry' supplies an explicit timing condition and 'ambiguous payments' identifies the target scenario, so an agent knows when to reach for this tool. It does not enumerate exclusions or name alternatives, so it is clear context rather than full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transaction_risk_scoreTransaction Risk ScoreARead-onlyIdempotentInspect
Estimate evidence-supported loss probability without selling insurance or assuming liability.; 0.004 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, and non-destructive behavior; the description adds deterministic, evidence-bounded, and signed behavior plus a cost figure. This is substantive context beyond the annotations, though the '0.004 USDC' cost is unlabeled, which prevents a higher score.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loads the main purpose, but the text is punctuated awkwardly ('liability.; 0.004 USDC.') and the cost fragment is unexplained. It is concise, though not cleanly structured enough for a 4.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the read-only annotations and the presence of an output schema, return values and side-effect safety need no elaboration. The main gaps are the unlabeled cost and the lack of guidance about how to supply the 'evidence' implied by the schema's open minProperties object.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero named parameters and 100% schema description coverage, the baseline is 4; there are no parameter fields for the description to clarify. The phrase 'evidence-supported' loosely signals what inputs may represent, but no explicit parameter documentation is needed under this schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening clause names a specific action ('Estimate') and a concrete deliverable ('evidence-supported loss probability'), and the qualifier 'without selling insurance or assuming liability' clarifies the tool is analytical rather than risk-bearing. It does not explicitly differentiate itself from the many risk-related sibling tools, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a use case—need a deterministic, evidence-supported loss-probability estimate—and explicitly excludes insurance/liability underwriting. However, it gives no guidance on when to use this tool instead of related siblings such as counterparty_risk_preflight or sla_risk_predictor, so the guidance remains mostly implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trust_anomaly_detectorTrust Anomaly DetectorARead-onlyIdempotentInspect
Detect material drift between comparable trust vectors without asserting misconduct.; 0.003 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, and non-destructive behavior, and the description goes well beyond them by adding cost ('0.003 USDC'), determinism, evidence-boundedness, signed output, and the explicit guardrail that it does not assert misconduct. These are meaningful behavioral disclosures beyond static annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely compact, front-loaded with the action, and every segment earns its place: purpose, cost, and behavioral guarantees. There is no redundant restatement of the tool name or title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema and annotations cover return values and safety, so the description does not need to repeat them. It supplies cost, determinism, and signing behavior, which are important invocation considerations. The main residual gap is that 'trust vector' is domain-specific and never defined, leaving some ambiguity about exact input expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero formal parameters, so the baseline is 4 per the rubric. The description adds context that inputs should be 'comparable trust vectors,' which is useful given the highly permissive schema. It does not define the internal shape of those vectors, but with no required parameters this is a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'Detect material drift between comparable trust vectors.' The qualifier 'without asserting misconduct' adds useful semantic precision. It does not explicitly distinguish the tool from sibling detectors, but the verb and object make the core purpose clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'comparable trust vectors' implies the intended use case, and the cost note gives an operational constraint. However, there is no explicit guidance about when to use this tool versus alternatives like signed_result_comparator or service_gap_detector, nor any when-not-to-use conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
try_serviceBInspect
Free representative preview, exact price and upgrade instructions for any PHION paid service; no wallet required.
| Name | Required | Description | Default |
|---|---|---|---|
| service | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does disclose two important traits: cost is zero ('Free') and no wallet/auth is needed. It stops short of other operational facts an agent needs — whether the preview is truncated/representative-only, rate limits, and whether calling it commits to anything.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence that front-loads the key benefit ('Free representative preview') before the outputs and the 'no wallet required' constraint. Nothing is wasted, though the clause stacking makes it slightly harder to parse at a glance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a 120+ value enum, no output schema, and no annotations, the description is adequate on the basics (free, no wallet, returns price and upgrade path) but does not explain what the 'representative preview' actually contains or how it relates to the real paid response, which matters given the huge service surface.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. 'For any PHION paid service' does convey that the single parameter selects a service, which is useful, but it adds no format, naming, or constraint detail beyond the 120+ value enum, and gives no hint about the enum itself.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('free representative preview') plus a concrete scope ('any PHION paid service') and the extra outputs (exact price, upgrade instructions). It is clearly distinguishable from a generic execute tool, but it never names or contrasts any of the many related siblings such as phion_get_service, phion_preflight, or phion_discover.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'No wallet required' implies the usage context (pre-payment evaluation, cold-start without credentials), so the when-to-use is inferable. However, it never states when NOT to use it (e.g. when a paid call is actually intended) nor points at the alternative tool to use in that case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verified_news_monitorCInspect
Literal query evidence across bounded public news sources; 0.006 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| source_urls | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must disclose behavior on its own. It does mention a cost of 0.006 USDC and the bounded news-source scope, which is useful, but it never explains what 'evidence' means, whether the tool is read-only, what happens when the query is not found, or if execution triggers a payment.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short, but it is a fragmented noun phrase rather than a structured tool definition. It omits essential information and reads more like a tagline than a functional specification, so the brevity is under-specification rather than conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With two undocumented parameters, no output schema, no annotations, and a large suite of sibling evidence/search tools, this description is far from adequate. An agent cannot know what to pass, what to expect back, or how this differs from the many comparable tools in the sibling list.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention query or source_urls at all. The agent receives no help understanding that source_urls is required, how many URLs to supply, what kind of URLs are accepted, or how query affects the behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The phrase 'Literal query evidence across bounded public news sources' suggests the tool returns evidence of a query restricted to news sources, but it lacks a clear verb such as 'fetch', 'verify', or 'search'. The purpose is vague enough that an agent could mistake it for a search tool or an evidence-certification tool, especially given many similar siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus siblings like source_backed_search, verified_web_extract, or multi_source_fact_bundle. The description gives no exclusions, prerequisites, or selection criteria, so an agent is left to guess based only on the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verified_web_extractCInspect
Bounded public web extraction with source, timestamp and hashes; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| source_urls | Yes | ||
| expected_fields | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It discloses that extraction is bounded, includes source/timestamp/hashes, and costs 0.003 USDC. However, it doesn't explain what 'verified' means behaviorally, whether it fetches live or cached content, or what happens on failure. The cost disclosure is useful but the verification semantics are vague.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence with useful specifics (bounded, source, timestamp, hashes, cost). It's front-loaded with the core action and scope. Could add a bit more context without becoming bloated, but it earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description is thin. It doesn't explain the return format, how verification works, what 'bounded' limits are (beyond maxItems 5 in schema), or how query and expected_fields shape the result. An agent would need to guess at critical invocation details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions 'source' which maps to source_urls, but doesn't explain query or expected_fields at all. The description adds minimal meaning beyond the schema's parameter names, leaving the agent to guess how query and expected_fields interact with extraction.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('extract') and resource ('public web') with bounded scope, and mentions source, timestamp, and hashes. It doesn't explicitly differentiate from siblings like source_backed_search or fetch_evidence, but the 'verified' and 'bounded' framing gives a clear sense of what it does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like source_backed_search, fetch_evidence, or verified_news_monitor. The description implies a use case (bounded public web extraction with verification) but doesn't state when to prefer it or what conditions make it appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verifyVerify a PHION ReceiptCRead-onlyIdempotentInspect
Free verification of a PHION signed receipt.
| Name | Required | Description | Default |
|---|---|---|---|
| content | No | ||
| receipt | Yes | ||
| sources | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false. The description adds only 'Free', which is a useful pricing trait not covered by annotations, but it omits what verification actually checks, what an invalid receipt produces, and any return behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is front-loaded and free of filler. However, its extreme brevity borders on under-specification for a tool with three parameters and no schema descriptions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and 0% schema description coverage, the description should explain return semantics and parameter roles. It does not, leaving an agent unclear about what verification produces and how to use the optional parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only hints at the required 'receipt' parameter via the phrase 'signed receipt'. It says nothing about the optional 'content' or 'sources' parameters, so it does not compensate for the undocumented schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('verification') and resource ('PHION signed receipt'), and adds the pricing trait 'Free'. It does not explicitly differentiate itself from siblings such as phion_preflight or phion_resolve, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, prerequisites, or alternatives. An agent must infer usage entirely from the tool name and title, with no exclusions or context about when this is preferable to sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
webhook_verifierWebhook VerifierBRead-onlyIdempotentInspect
Fail-closed RFC 9421-style webhook preflight requiring trusted key, algorithm, nonce, freshness, replay status and signature coverage of method, target and content; 0.003 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| event | Yes | ||
| verification | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish it as read-only, idempotent, non-destructive and closed-world, so the safety profile is covered. Beyond that, the description adds real behavioral signal: it is 'fail-closed,' enumerates the checks it enforces (trusted key, algorithm, nonce, freshness, replay status, signature coverage), and discloses a cost of '0.003 USDC.' This is meaningful context the annotations do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the core concept ('Fail-closed RFC 9421-style webhook preflight') before enumerating requirements and cost. Every clause carries information, though the packing of so many checks makes it slightly hard to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two nested objects, 0% schema descriptions and no output schema, the description supplies a fair amount (the enforcement model, the checked properties, the cost). But it omits any explanation of the `event` payload shape and does not state what the caller receives on pass vs. fail, which matters for a fail-closed gate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for two nested objects, so the description must carry parameter meaning. It usefully enumerates the verification requirements (trusted key, algorithm, nonce, freshness, replay, coverage of method/target/content), which maps onto the `verification` sub-fields. However, it says nothing about the `event` object (delivery_id, timestamp, payload) or how freshness ties to the timestamp, leaving half the parameter surface unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: a 'webhook preflight' that verifies signatures 'RFC 9421-style.' It is clearly a verification/validation gate for webhooks, distinguishing it functionally from the many payment and policy preflights among its siblings. It stops short of explicitly naming which sibling to prefer, but the purpose itself is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the word 'preflight' – the reader can infer it runs before accepting a webhook. There is no statement of when to use it versus the adjacent preflight/verify siblings or any precondition (e.g., call before processing the payload). No exclusions or alternatives are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
x402_quote_comparatorx402 Quote ComparatorCRead-onlyIdempotentInspect
Validate x402 v2 fields and compare only quotes sharing asset and decimals; 0.002 USDC.
| Name | Required | Description | Default |
|---|---|---|---|
| quotes | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds that x402 v2 fields are validated and only same-asset/same-decimals quotes are compared, but it omits failure behavior, output shape, and what '0.002 USDC' represents.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence is concise and front-loads the validate/compare actions. However, the unexplained '; 0.002 USDC' fragment adds confusion rather than clarity, and the description is too terse for a nested 64-item schema. It is not verbose, but structure is weakened by the stray fragment.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The input is high complexity: nested quote objects, five required fields per quote, and min/max item bounds, with no output schema and annotations only covering safety. The description does not explain validation rules, comparison output, error cases, or the 0.002 USDC figure. An agent lacks enough context to invoke this confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter is a nested array with five required fields, minItems 2, and maxItems 64. The description generically mentions asset, decimals, and x402 v2 fields but leaves network, payTo, amount, item shape, requiredness, and bounds undocumented. It does not compensate for the absent schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States specific verbs 'Validate' and 'compare' with the resource 'quotes'. It constrains comparison to quotes sharing asset and decimals, but does not differentiate this tool from sibling comparators or payment tools. The trailing '0.002 USDC' is unexplained and does not clarify purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use, when-not, or alternative tool is named. The only guidance is an internal correctness constraint ('only quotes sharing asset and decimals'), not selection guidance against sibling tools such as signed_result_comparator or payment_route_selector.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
x402_v2_routerx402 v2 RouterCRead-onlyIdempotentInspect
Select an exactly compatible x402 v2 payment requirement without receiving private keys.; 0.003 USDC. Deterministic, evidence-bounded and signed.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry readOnlyHint, idempotentHint, and destructiveHint, and the description adds value beyond them: determinism, signed output, evidence-bounded behavior, and the privacy property of not receiving private keys. These traits are not present in any structured field and meaningfully shape expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
At roughly twenty words it is appropriately brief and leads with the primary action. However, the semicolon-separated fragment '0.003 USDC.' and the awkward 'without receiving private keys' clause give it a stitched-together feel rather than clear structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Annotations and the output schema cover safety and return shape, and the description adds determinism and signing behavior. Still, an agent cannot determine what payload to send or what '0.003 USDC' refers to, which is a real gap for a tool whose input schema is intentionally opaque.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema exposes zero named parameters yet requires at least one property (minProperties: 1, additionalProperties: true), and the description never explains what to pass. Although the 0-param baseline is 4, this schema is not truly parameter-free — it accepts an arbitrary object — and the description fails to compensate for that ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a verb and resource ('Select an exactly compatible x402 v2 payment requirement'), which distinguishes it from sibling tools like x402_quote_comparator. However, the purpose is muddied by unexplained fragments ('0.003 USDC', 'evidence-bounded') that appear disconnected from the primary action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as x402_quote_comparator or payment_preflight. No conditions, exclusions, or prerequisites are stated, so an agent must infer the invocation context from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
- Added
phion_enrich_basic - Changed
try_service1 field changed- changed
Input schema / properties / service / enumPrevious value: -[ - "attest", - "payment-firewall", - "web-evidence", - "mandate-reserve", - "agent-inspection", - "commerce-journey", - "transaction-assurance", - "transaction-recovery", - "rwa-asset-due-diligence", - "rwa-compliance", - "rwa-nav-reserve", - "rwa-transaction-assurance", - "rwa-corporate-actions", - "schema-normalize", - "sanctions-screening-evidence", - "agent-reputation-evidence", - "counterparty-risk-preflight", - "tool-output-firewall", - "delegation-scope-guard", - "memory-write-guard", - "mcp-manifest-firewall", - "tool-call-policy-guard", - "data-egress-preflight", - "agent-budget-guard", - "idempotency-replay-guard", - "human-approval-policy", - "secret-redaction-preflight", - "oauth-token-audience-guard", - "redirect-callback-validator", - "tool-capability-drift-monitor", - "mcp-server-identity-evidence", - "agent-task-handoff-receipt", - "context-provenance-labeler", - "signed-result-comparator", - "service-sla-attestation", - "payment-route-selector", - "x402-quote-comparator", - "payment-receipt-reconciler", - "duplicate-charge-detector", - "subscription-spend-guard", - "webhook-verifier", - "delivery-evidence", - "inter-agent-policy-evaluator", - "concurrency-guard", - "resource-cycle-guard", - "abandoned-tool-detector", - "manifest-version-diff", - "dependency-provenance-assessment", - "conflict-resolution", - "data-freshness-certificate", - "domain-ownership-evidence", - "purpose-bound-consent", - "retention-deletion-receipt", - "interrupted-task-recovery", - "portable-observability-audit", - "index-feed", - "verified-web-extract", - "entity-enrichment-evidence", - "social-source-evidence", - "market-data-snapshot", - "onchain-evidence", - "verified-news-monitor", - "multi-source-fact-bundle", - "document-to-verified-json", - "source-backed-search", - "live-data-freshness", - "person-enrichment-evidence", - "company-enrichment-evidence", - "contact-enrichment-evidence", - "mcp-2026-compatibility-gateway", - "durable-agent-task", - "task-checkpoint-evidence", - "task-cancel-assurance", - "agent-approval-relay", - "capability-negotiation-preflight", - "mcp-catalog-cache-guard", - "oauth-issuer-binding-evidence", - "mcp-a2a-task-bridge", - "quote-freshness-guard", - "delegated-credential-guard", - "agent-session-continuity", - "task-lease-guard", - "tool-result-schema-validator", - "agent-memory-provenance", - "payment-delivery-atomicity", - "service-failover-selector", - "agent-rate-limit-negotiator", - "execution-cost-estimator", - "cross-agent-receipt-bundle", - "capability-verification", - "capability-benchmark", - "sybil-reputation-guard", - "agent-behavior-fingerprint", - "trust-anomaly-detector", - "economic-loop-detector", - "provider-quality-predictor", - "transaction-risk-score", - "payment-optimizer", - "x402-v2-router", - "erc8004-identity-evidence", - "erc8004-reputation-intelligence", - "delegated-spend-policy", - "ap2-mandate-bridge", - "a2a-transaction-bridge", - "mcp-transaction-gateway", - "proof-of-service", - "automated-dispute-bundle", - "transaction-recovery-v2", - "reputation-update-receipt", - "agent-economic-graph", - "market-demand-predictor", - "service-gap-detector", - "price-discovery-engine", - "sla-risk-predictor", - "dynamic-service-pricing", - "autonomous-procurement", - "multi-agent-escrow", - "cross-protocol-receipt", - "phion-inference", - "phion-search", - "phion-enrich", - "phion-intelligence" -]New value: +[ + "attest", + "payment-firewall", + "web-evidence", + "mandate-reserve", + "agent-inspection", + "commerce-journey", + "transaction-assurance", + "transaction-recovery", + "rwa-asset-due-diligence", + "rwa-compliance", + "rwa-nav-reserve", + "rwa-transaction-assurance", + "rwa-corporate-actions", + "schema-normalize", + "sanctions-screening-evidence", + "agent-reputation-evidence", + "counterparty-risk-preflight", + "tool-output-firewall", + "delegation-scope-guard", + "memory-write-guard", + "mcp-manifest-firewall", + "tool-call-policy-guard", + "data-egress-preflight", + "agent-budget-guard", + "idempotency-replay-guard", + "human-approval-policy", + "secret-redaction-preflight", + "oauth-token-audience-guard", + "redirect-callback-validator", + "tool-capability-drift-monitor", + "mcp-server-identity-evidence", + "agent-task-handoff-receipt", + "context-provenance-labeler", + "signed-result-comparator", + "service-sla-attestation", + "payment-route-selector", + "x402-quote-comparator", + "payment-receipt-reconciler", + "duplicate-charge-detector", + "subscription-spend-guard", + "webhook-verifier", + "delivery-evidence", + "inter-agent-policy-evaluator", + "concurrency-guard", + "resource-cycle-guard", + "abandoned-tool-detector", + "manifest-version-diff", + "dependency-provenance-assessment", + "conflict-resolution", + "data-freshness-certificate", + "domain-ownership-evidence", + "purpose-bound-consent", + "retention-deletion-receipt", + "interrupted-task-recovery", + "portable-observability-audit", + "index-feed", + "verified-web-extract", + "entity-enrichment-evidence", + "social-source-evidence", + "market-data-snapshot", + "onchain-evidence", + "verified-news-monitor", + "multi-source-fact-bundle", + "document-to-verified-json", + "source-backed-search", + "live-data-freshness", + "person-enrichment-evidence", + "company-enrichment-evidence", + "contact-enrichment-evidence", + "mcp-2026-compatibility-gateway", + "durable-agent-task", + "task-checkpoint-evidence", + "task-cancel-assurance", + "agent-approval-relay", + "capability-negotiation-preflight", + "mcp-catalog-cache-guard", + "oauth-issuer-binding-evidence", + "mcp-a2a-task-bridge", + "quote-freshness-guard", + "delegated-credential-guard", + "agent-session-continuity", + "task-lease-guard", + "tool-result-schema-validator", + "agent-memory-provenance", + "payment-delivery-atomicity", + "service-failover-selector", + "agent-rate-limit-negotiator", + "execution-cost-estimator", + "cross-agent-receipt-bundle", + "capability-verification", + "capability-benchmark", + "sybil-reputation-guard", + "agent-behavior-fingerprint", + "trust-anomaly-detector", + "economic-loop-detector", + "provider-quality-predictor", + "transaction-risk-score", + "payment-optimizer", + "x402-v2-router", + "erc8004-identity-evidence", + "erc8004-reputation-intelligence", + "delegated-spend-policy", + "ap2-mandate-bridge", + "a2a-transaction-bridge", + "mcp-transaction-gateway", + "proof-of-service", + "automated-dispute-bundle", + "transaction-recovery-v2", + "reputation-update-receipt", + "agent-economic-graph", + "market-demand-predictor", + "service-gap-detector", + "price-discovery-engine", + "sla-risk-predictor", + "dynamic-service-pricing", + "autonomous-procurement", + "multi-agent-escrow", + "cross-protocol-receipt", + "phion-inference", + "phion-search", + "phion-enrich-basic", + "phion-enrich", + "phion-intelligence" +]
1 tool update
- Changed
payment_capability_preflight1 field changed- added
Input schema / properties / evm_account_kindAdded value: +{ + "enum": [ + "standard_eoa", + "delegated_eip7702", + "smart_contract", + "unknown" + ] +}
Related MCP Connectors
Free discovery and preflight for trusted evidence-backed AIOS agent services.
Discover agent services and buy IssueFoundry's fractional-cent x402 reliability utilities.
Free AI crawler checks, Agent Surface audits, A2A actions and optional x402 evidence packs.
Discover and price paid agent services across the live x402 network, by three.ws.
Related MCP Servers
AlicenseBqualityBmaintenancex402 Discovery Hub. Search engine for the agent economy with 1450+ services indexed. Search by keyword, browse by category, free service submission.1248 npm4MIT- FlicenseNot gradedqualityAmaintenanceEnables agents to discover and verify paid agent services through x402 payment validation and release gate checks.-

AgentBodega MCPofficial
AlicenseAqualityFmaintenanceEnables agents to search and inspect live service offerings, generate x402 payment snippets, and understand blockchain-only balance policies.458 npm3MIT- AlicenseAqualityCmaintenanceEnables agents to verify claims against live web evidence with calibrated confidence, paying per call via x402 and receiving offline-verifiable signed receipts.3MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.
social_source_evidenceDInspect
Attributable public social-source observations with identity limitations; 0.006 USDC.
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full responsibility for behavioral disclosure. It does mention 'identity limitations' and a cost of 0.006 USDC, but it does not state whether the operation mutates state, requires authentication, has rate limits, or what outcomes an agent should expect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The text is short but reads as a noun-phrase fragment rather than a usable definition. It front-loads no actionable information, and the cost/limitation details are appended without explanation, so brevity comes at the expense of substance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given three parameters, a nested object, no output schema, and dozens of sibling tools, this description is far too sparse. An agent cannot determine how to construct a valid request, what response to expect, or how this tool fits among its alternatives.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not mention query, subject, or source_urls. The required source_urls parameter is undocumented, and the nested subject object is entirely unexplained, so the agent cannot infer parameter semantics from either source.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the tool's domain (public social-source observations) and mentions attribution and identity limitations, but it lacks a verb stating what the tool actually does – e.g., fetch, verify, or record evidence. Without that, an agent cannot determine how it differs from siblings like fetch_evidence, source_backed_search, or onchain_evidence.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool rather than the many evidence-related siblings. The only contextual clue is the 0.006 USDC price, which implies a paid operation, but no selection criteria, prerequisites, or exclusion conditions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.