AgentWatch
Server Details
Read-only watchtower for AI agents on-chain: decoded receipts, plan vs execution, Safe audits.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-03-26
- URL
TDQS
Scored across 21 tools
Most tools have clear, distinct purposes, but the seven 'get_' analytics tools (get_benchmark, get_capabilities, get_fidelity, get_feed, get_tca, get_leaderboard, get_dashboard) could be confused by an agent unsure which metric to request. The descriptions provide specific use-case guidance, reducing ambiguity.
Nearly all tools follow a consistent verb_noun snake_case pattern (e.g., add_watch, get_feed, list_watched, set_recovery_owner). The exception is watch_demo_set, which uses a noun_verb_noun structure, slightly breaking the pattern.
With 21 tools, the server is on the heavy side and some consolidation is possible (e.g., get_dashboard overlaps with get_alerts and get_feed). However, each tool targets a distinct capability, so it is borderline rather than excessive.
Core read and analytics operations are well covered, but a basic lifecycle operation is missing: there is no remove_watch to stop watching an address. Additionally, no tool lists or retrieves intent status/history, leaving gaps for agents needing to manage watches or intents.
Available Tools
21 toolsadd_watchBInspect
Start watching an address. Use when onboarding a new agent or Safe.
| Name | Required | Description | Default |
|---|---|---|---|
| label | No | Optional short label (e.g. Agent EOA) | |
| chains | No | ||
| address | Yes | Address to watch (0x…40 hex) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and largely drops it. It never states whether the watch persists, whether it is idempotent for an already-watched address, whether watched data goes stale and needs refresh_watches, or what permission/ownership is required. 'Start watching' implies ongoing monitoring but nothing about lifecycle or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler, and the action statement is front-loaded before the usage hint. Nothing is repeated from the schema or title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter mutating tool with no annotations and no output schema, the description is too thin: no statement of effect, persistence, idempotency, or error behavior, and one parameter format is unexplained. An agent could invoke it, but not confidently or safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% (address and label are documented, chains is bare), and the description adds no parameter meaning at all beyond the word 'address'. The undocumented 'chains' string parameter — whose format is critical for scoping the watch — is left unexplained in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Start watching an address'), which is unmistakable and distinct from read-oriented siblings like list_watched or get_dashboard. It does not, however, contrast itself with related siblings such as refresh_watches or watch_demo_set, so an agent still has to infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit triggering context ('Use when onboarding a new agent or Safe'), which tells the agent when this tool belongs in a workflow. It gives no when-not guidance and names no alternative tool for adjacent cases like refreshing or removing watches.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_agent_passportAInspect
Read the declared identity of a watched agent (name, model provider, how often it runs, role, purpose, operator) and whether its public passport page is published.
| Name | Required | Description | Default |
|---|---|---|---|
| address | Yes | Watched agent address (0x…40 hex) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It implies a safe read and discloses one interesting state ('whether its public passport page is published'), but says nothing about permission requirements, behavior when the address is not watched, or error/empty-result handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with no filler; the parenthetical field list is dense but directly useful. Nothing is wasted, though the list could have been tightened slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the returned fields, which is exactly what the agent needs to decide whether to call it. It stops short of covering failure/precondition behavior for addresses that are not watched.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single 'address' parameter, so the schema already documents the 0x…40 hex format. The description adds only the qualifier 'watched', which slightly narrows the valid input but no format or resolution detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') plus the exact resource ('declared identity of a watched agent') and enumerates the fields returned (name, model provider, run frequency, role, purpose, operator, passport publication status). This cleanly separates it from the write-side sibling set_agent_passport without needing to name it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the phrase 'watched agent' hints that the address must already be under watch (add_watch/list_watched), but there is no explicit when-to-use statement, no mention of the set_agent_passport counterpart, and no guidance on what happens for an unwatched address.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_alertsBInspect
List alerts scoped to this principal's watched addresses (empty watch list → call watch_demo_set first). Use when checking what needs attention.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It usefully discloses scope (only watched addresses) and a dependency on the watch list being populated, but says nothing about permissions, read-only safety, ordering, or result volume for a listing operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the scope and followed by the usage cue and prerequisite. Efficient, though the 'Use when checking what needs attention' clause is mildly generic.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read tool with no output schema or annotations, the description covers scope and the watch-list prerequisite but omits any explanation of the 'status' filter, leaving a real gap in how to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The sole parameter 'status' has 0% schema description coverage and is never mentioned in the description, so the agent has no guidance on what values are valid or what filtering it performs. The description fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List') and resource ('alerts') with an explicit scope ('scoped to this principal's watched addresses'). Sibling tools like list_watched and watch_demo_set are distinguishable, though the description does not name any alternative directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a clear usage context ('Use when checking what needs attention') and a concrete prerequisite that routes to a sibling ('empty watch list → call watch_demo_set first'). It lacks any explicit when-not-to-use or a named alternative list tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_benchmarkBInspect
Counterfactual P&L / agent alpha vs do-nothing baselines. Use for performance honesty checks. Check matched_agent — false means the key has no indexed activity.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_key | Yes | Agent EOA address (0x…40 hex) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses an output interpretation ("matched_agent — false means the key has no indexed activity"), which is valuable given there is no output schema. However, it omits auth requirements, cost/compute expectations, and time-window behavior for a metrics tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no filler; the core purpose leads. It is dense with domain jargon ("counterfactual P&L", "agent alpha") that slightly reduces immediate clarity, keeping it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description partially compensates by explaining the matched_agent flag, but leaves the return shape, metric units, and observation window undefined for a benchmark tool. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter with 100% schema description coverage, so the schema already documents agent_key as an EOA address. The description's reference to "the key" adds no format or constraint detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific, non-obvious purpose: "Counterfactual P&L / agent alpha vs do-nothing baselines," which distinguishes it from generic siblings like get_fidelity or get_leaderboard by naming the counterfactual baseline concept. The verb is implied ("get") and no sibling is named explicitly, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use for performance honesty checks" gives an implied context for when to reach for this tool, but names no alternatives and provides no when-not guidance. An agent gets a rough sense of intent but must still guess how it differs from get_fidelity or run_audit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_capabilitiesBInspect
Blast-radius / historical capability for an agent key. Use before trusting an agent with funds.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_key | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a read of historical capability data but discloses nothing about permissions, rate limits, data freshness, or what the response contains. Only the safety-intent framing is offered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two terse sentences, front-loaded with the resource and followed by the usage directive. Nothing is wasteful, though the terse fragment style borders on under-specification rather than true economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-parameter read tool with no output schema and no annotations, the description conveys purpose and timing but leaves the meaning of 'blast-radius' and the returned data opaque. Adequate but with clear gaps an agent would notice.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there is one parameter. The description only implies 'agent key' maps to agent_key; it adds no format, source, or validation meaning beyond the schema. With low coverage it should compensate more than it does.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource ('historical capability for an agent key') and introduces a scope concept ('blast-radius'), so an agent can grasp the intent. It does not explicitly differentiate from close siblings like get_agent_passport, but the resource is named clearly enough to distinguish.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete when-to-use directive: 'Use before trusting an agent with funds.' That is a real usage context. It names no alternatives or when-not-to-use conditions, which keeps it short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dashboardAInspect
Persona-tailored dashboard brief (fidelity, alerts, pending plans, recent planned-vs-executed). NEW USERS with empty watches: call watch_demo_set first, then get_dashboard again. Narrate in plain language; ask the user to Always allow the AgentWatch MCP App widget when the host prompts.
| Name | Required | Description | Default |
|---|---|---|---|
| persona | No | Optional override; otherwise uses saved persona |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose non-obvious behavior: the empty-watch state and its remediation path, plus a host permission prompt for the MCP App widget. It does not state read-only semantics or error behavior, leaving a modest gap for a read operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, purpose front-loaded, with the prerequisite and host-prompt guidance following. Every sentence earns its place, though the narration/permission instruction is slightly tacked-on relative to the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-parameter read tool with no output schema, the description covers what the brief contains and the new-user prerequisite. It omits the read-only nature and any error/empty state beyond the empty-watch case, but is otherwise sufficient to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single 'persona' parameter is already documented as an optional override of the saved persona. The description only echoes this with 'persona-tailored,' adding no syntax, format, or fallback detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('dashboard brief') and enumerates its contents (fidelity, alerts, pending plans, planned-vs-executed), which implicitly positions it as an aggregate over siblings like get_alerts and get_fidelity. It does not explicitly name those siblings, so the differentiation is inferential rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives one concrete precondition and workflow: new users with empty watches must call watch_demo_set first, then re-call get_dashboard. It does not state when to prefer this over the granular sibling tools, so it stops short of full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_feedBInspect
Decoded activity feed for an address. Use when reviewing what an agent did on-chain.
| Name | Required | Description | Default |
|---|---|---|---|
| since | No | Unix seconds lower bound | |
| address | Yes | Watched address (0x…40 hex) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a read-only feed but does not disclose permissions, rate limits, pagination, or the distinction between decoded and raw data beyond the word 'decoded.'
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with purpose followed by usage; no filler. The fragment 'Decoded activity feed for an address' is efficient, though 'decoded' could use a brief qualifier without wasting space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description omits return format, pagination behavior, default time range for 'since', and what constitutes a 'decoded' activity entry. These are important for an agent to invoke and interpret the feed correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both 'address' and 'since' fully. The description adds no parameter syntax or format details beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource ('Decoded activity feed for an address') and adds a usage context ('reviewing what an agent did on-chain'). However, it does not explicitly differentiate itself from sibling read tools like get_alerts or get_dashboard, which could also surface on-chain activity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear when-to-use condition: reviewing what an agent did on-chain. It does not mention when not to use it or name any alternative tools, leaving the agent to infer that this is the primary feed tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_fidelityBInspect
Intent↔execution fidelity score for an agent or Safe. Use for accountability / coverage meter. Check matched_agent — false means the address has no indexed activity (as signer or as the Safe that executed), so the score is vacuous.
| Name | Required | Description | Default |
|---|---|---|---|
| agent_key | Yes | Agent EOA address (0x…40 hex) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses an edge case (matched_agent=false makes the score vacuous), but omits other traits such as whether it is read-only, auth requirements, or cost/rate behavior, leaving meaningful gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with purpose then usage then a caveat, with little waste. The final clause is dense and references matched_agent without establishing it as a return field, slightly hurting readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description only partly compensates by flagging the matched_agent field. The score's scale/range, other return fields, and the surrounding context of the fidelity metric are left undefined.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema already documents the single agent_key parameter, setting the baseline at 3. The phrase "for an agent or Safe" hints the key may accept either address type beyond the schema's "Agent EOA address", but it is vague and not an explicit parameter clarification.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource and computation ("Intent↔execution fidelity score for an agent or Safe"), so an agent knows exactly what is being returned. It does not, however, distinguish itself from sibling metric tools like get_benchmark or get_tca, which are also read-only scores.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use for accountability / coverage meter" gives implied usage context, but there is no explicit when-to-use vs. when-not guidance and no alternative sibling is named. The agent must infer selection from the purpose statement alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_leaderboardAInspect
Public standings for agents with a published passport, ranked from receipts (declared-first coverage, fidelity, reliability, hygiene, sustained work). Pass agent_key to get that agent's own rank, its component scores, and what it would need to climb. Read-only: nothing here can be self-reported.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many ranked rows to return (default 10, max 50) | |
| agent_key | No | Optional: address to locate in the standings |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningful work: it declares 'Read-only' and discloses the trust model ('nothing here can be self-reported'), telling the agent the data is derived from receipts rather than user input. It omits rate limits and pagination behavior, but the core behavioral profile is covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the core purpose and scope front-loaded, followed by the agent_key behavior and the read-only caveat. The parenthetical ranking list is dense but earns its place by defining the score's composition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must describe returns, and it does so adequately: ranked rows by default, and per-agent rank plus component scores and improvement guidance when agent_key is passed. The limit parameter's default and cap live in the schema, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description goes beyond the schema by explaining what agent_key actually returns (that agent's own rank, component scores, and what it would need to climb), which the schema's 'address to locate in the standings' does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (public standings) with a precise scope (agents with a published passport) and enumerates the ranking inputs (coverage, fidelity, reliability, hygiene, sustained work). It is distinguishable from get_agent_passport and get_fidelity in practice, though it never names a sibling to route against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage via 'Pass agent_key to get that agent's own rank,' which tells the agent when to supply the optional parameter. However, it never states when to prefer this over get_agent_passport or get_fidelity, nor any exclusions, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_pending_proposalsBInspect
Pending Safe multisig proposals from the Safe Transaction Service (indexer confidence). Pass safe=0x… or omit to scan all watched addresses.
| Name | Required | Description | Default |
|---|---|---|---|
| safe | No | Optional Safe address; default = all watched | |
| chains | No | Optional comma-separated chain ids |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. The '(indexer confidence)' note hints that results are indexer-derived and possibly incomplete, which is genuinely useful, but there is no mention of return shape, freshness, pagination, or permissions for a read tool with no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the resource front-loaded and the scoping rule second; every clause carries information. The parenthetical '(indexer confidence)' is slightly vague but still earns its place as a data-quality cue.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-required-param, no-annotation, no-output-schema tool, the definition covers what is fetched and how to scope it, but a reader cannot infer whether results are cached, how fresh they are, or what the response contains beyond 'pending proposals'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both parameters are already documented in the schema, and the description largely restates the same 'omit = all watched' default. It adds only a weak format hint ('0x…') and nothing about the 'chains' parameter's expected values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Pending Safe multisig proposals from the Safe Transaction Service'), so the agent knows exactly what object is returned. No sibling covers proposal retrieval, so differentiation is implicitly satisfied, but the description never names or contrasts an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It tells the agent how to scope the call ('Pass safe=0x… or omit to scan all watched addresses'), which is a usable condition, but gives no when-to-use vs. when-not guidance relative to siblings like get_feed, get_alerts, or get_dashboard.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_safe_defaultsAInspect
Return the default Safe ownership plan for an agent EOA: threshold 1 + agent + selected recovery (1/2 when one recovery is set). Call before register_intent for Safe deploys. AgentWatch never deploys or signs.
| Name | Required | Description | Default |
|---|---|---|---|
| chain | No | Optional chain hint for the intent constraints | |
| purpose | No | Optional purpose filter (e.g. trading) | |
| agent_key | Yes | Agent EOA that will be an owner/signer | |
| all_recoveries | No | If true, include all matching recoveries (1/n) instead of just the default one | |
| recovery_label | No | Optional label to pick among several recoveries |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and delivers the key trait: 'AgentWatch never deploys or signs,' telling the agent this is a non-mutating planning/read operation. It also conveys that the result depends on whether a recovery is set. It does not mention auth requirements, rate limits, or whether results are cached, so it is strong but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with zero filler, front-loaded with what is returned before the workflow instruction and the safety guarantee. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by describing the shape of the returned plan and the rule that determines it, plus the non-deploy/non-sign guarantee. For a 5-parameter, single-required-parameter read tool this is nearly complete, with only return-format/pagination-style details absent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes slightly beyond the schema by explaining the default-selection logic ('selected recovery', '1/2 when one recovery is set'), which clarifies what all_recoveries and recovery_label actually change. chain, purpose and agent_key remain explained only by the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource — 'Return the default Safe ownership plan for an agent EOA' — and even spells out the returned plan (threshold 1 + agent + selected recovery, 1/2 when one recovery is set). It is not a tautology and is clearly distinct in intent from siblings like list_recovery_owners or register_intent. It lacks an explicit 'this is not X' contrast, but the specificity alone makes the tool identifiable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Call before register_intent for Safe deploys' gives explicit sequencing and names the related downstream tool, so an agent knows when in a workflow to use it. It does not state when to skip it or what happens if called standalone, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_tcaBInspect
TCA market-context snapshot for an event_id (chain:txHash). Use when judging execution quality.
| Name | Required | Description | Default |
|---|---|---|---|
| event_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden and largely fails to: 'snapshot' weakly implies a read, but there is no word on auth requirements, rate limits, caching/freshness, error behavior, or whether anything is mutated. For a tool with zero structured behavior hints, this is a substantial gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the resource and scope, with nothing obviously wasted. It is efficient if terse; the terseness is arguably under-specification rather than over-verbosity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description should compensate, but 'market-context snapshot' gives no indication of what is returned (spread, slippage, liquidity, benchmark deltas). Combined with unexplained 'TCA' and an unexplained event_id provenance, an agent cannot confidently call this without guessing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the schema only says event_id is a string, so the description does add real value by specifying the expected format '(chain:txHash)'. However it does not say where an event_id comes from, its uniqueness/validity constraints, or what happens on an unknown ID.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a resource (TCA market-context snapshot) and its scope (per event_id), which is more than a restated name, but the verb is left implicit and 'TCA' is unexplained jargon an agent may not resolve. It also never differentiates itself from plausibly related siblings such as get_fidelity or get_benchmark, so the purpose is only partly pinned down.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use when judging execution quality' gives an explicit use condition rather than leaving usage implied. It names no alternatives and states no exclusions, which caps it below a 5 (unlike the sibling-routing in the high-quality calibration example), but the context is clear enough to act on.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_recovery_ownersAInspect
List this user's saved recovery wallets (human co-signers for default 1/2 Safes). Each may have a label/purpose — users often keep several for different vaults.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full behavioral burden. It conveys valuable entity semantics — these are human co-signers tied to 1/2 Safes and may carry labels/purposes — but says nothing about auth requirements, pagination, ordering, or whether an empty result is possible. Read-only is only implied by the verb 'List'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the core action front-loaded and the clarifying domain detail second. Every clause earns its place; nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read tool with no output schema, the description gives enough semantic grounding to know what the returned entries represent. Minor gaps remain — result shape, ordering, and pagination are unaddressed — but nothing critical to invoking it is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies. The description correctly implies no filtering inputs are needed, though it does not explicitly state that results are scoped automatically to 'this user' (it implies it).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (this user's saved recovery wallets), and adds domain context that disambiguates it — human co-signers for default 1/2 Safes, with optional labels. The verb alone separates it from siblings set_recovery_owner and remove_recovery_owner, so an agent can pick it without opening another schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use statement, no prerequisites, and no named alternatives. The remark that 'users often keep several for different vaults' hints at why the data matters but does not tell the agent when to call this versus set_recovery_owner or get_safe_defaults.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_watchedAInspect
List watched addresses for this connector principal. NEW USERS: if empty, call watch_demo_set next (do not invent a global feed), then get_dashboard.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose one real behavioral trait: the list can be empty and what the empty state implies. However, it says nothing about permissions, per-principal auth requirements, or whether the returned set is cached/stale, which matters for a watch-management surface.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, no padding, and the primary action is front-loaded ahead of the empty-state guidance. Every sentence carries actionable content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, non-destructive list call with no output schema, the description covers the essentials: what is listed, the scope, and the empty-state next steps. It leaves out only the shape/ordering of the returned list, which the missing output schema would otherwise have to carry.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the schema is fully closed (additionalProperties: false), so there is no parameter semantics to add; the baseline of 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('watched addresses') with scope ('for this connector principal'), which cleanly separates it from add_watch, refresh_watches, and watch_demo_set. It doesn't explicitly contrast itself with siblings, but the scoping phrase does most of the differentiating work.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete conditional workflow: on an empty result, call watch_demo_set then get_dashboard, and explicitly warns against inventing a global feed. It stops short of stating when to prefer this listing over refresh_watches or get_feed, so it is strong but not fully exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
refresh_watchesAInspect
Force an immediate poller pass over this principal's watched addresses (catch up txs from explorers into AgentWatch). Use after add_watch / a fresh broadcast when get_feed is still empty. Optional address / chains to focus the pass.
| Name | Required | Description | Default |
|---|---|---|---|
| chains | No | Optional csv of chains (default: ethereum,base first, then the rest) | |
| address | No | Optional single address to refresh (0x…40 hex) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It discloses the mechanism (poller pass writing explorer txs into AgentWatch), which is useful, but says nothing about latency, whether the pass is synchronous or blocking, cost/rate sensitivity, or whether it can be safely repeated. Adequate but incomplete for a trigger-style operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and effect, then usage guidance, then parameter notes. No filler, though the last sentence is somewhat redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-output-schema, no-annotation trigger tool, the description leaves open what the caller observes after the pass (does it return results, or must get_feed be re-checked?) and whether the refresh blocks. The core behavior is covered, but these operational gaps matter for correct use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so both optional parameters are already documented in the schema, including the chain default. The description only restates their purpose ('focus the pass') without adding new semantics, so the baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('force an immediate poller pass'), its scope ('this principal's watched addresses'), and its effect ('catch up txs from explorers into AgentWatch'). This is clearly distinguishable from siblings like add_watch (registration) and get_feed (read-only retrieval).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the trigger conditions ('after add_watch / a fresh broadcast when get_feed is still empty') and references the sibling get_feed as the fallback state. It stops short of stating when NOT to use it (e.g., during normal operation or high-frequency polling), which keeps it from a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
register_intentAInspect
Declare an intent BEFORE broadcasting. Returns {intent_id, pairing_code}. Append the 16-hex pairing_code as the last 8 bytes of calldata (no magic prefix), then broadcast — EXCEPT bare ETH transfers to contracts (empty data + value>0): appending reverts on Safe/fallback handlers; register without a suffix and use expected.steps for multi-leg flows (e.g. approve+supply). Pairing codes are single-use and only consumed by successful (non-reverted) txs. Use deadline_minutes (relative) — normalized to absolute constraints.deadline at registration.
| Name | Required | Description | Default |
|---|---|---|---|
| expected | No | Structured shape. Use steps: ['approve','other'] (or similar) so one intent covers approve+supply/swap companions. | |
| agent_key | Yes | ||
| declared_at | No | ||
| description | Yes | ||
| deadline_minutes | No | Relative window; stored as absolute constraints.deadline |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden and delivers real behavioral detail: the returned fields, that pairing codes are single-use and only consumed by non-reverted txs, and that appending reverts on Safe/fallback handlers for bare ETH transfers. It does not state auth/permission requirements or whether registration is on- or off-chain, leaving some gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core imperative, and every sentence carries operational weight (return shape, suffix rule, exception, single-use semantics, deadline normalization). Dense em-dash clauses make it slightly harder to parse than it needs to be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex EVM-calldata tool with 5 params at 40% coverage, no annotations, and no output schema, the description covers the critical mechanics (return shape, suffix placement, exception path, multi-leg handling) well. It is not fully complete given the undocumented parameters and unstated auth model, but nothing essential for invoking it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40%, so the description must compensate. It adds real meaning for deadline_minutes (relative, normalized to absolute constraints.deadline) and expected.steps, but leaves agent_key and declared_at unexplained. Partial compensation justifies the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Declare an intent') plus the exact workflow ('BEFORE broadcasting') and what it returns ({intent_id, pairing_code}). No sibling tool does anything similar (all are get/list/set/audit tools), so it is trivially distinguishable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use ('Declare an intent BEFORE broadcasting') and an explicit when-not exception for bare ETH transfers to contracts (empty data + value>0) that would revert, plus redirection to expected.steps for multi-leg flows. No named alternative tool is offered, but none exists among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_recovery_ownerCInspect
Remove a saved recovery address for this user.
| Name | Required | Description | Default |
|---|---|---|---|
| address | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'Remove' implies a destructive, likely irreversible mutation, yet the description says nothing about required permissions, confirmation steps, or the effect on account recoverability when the last address is removed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short sentence with zero filler and the action front-loaded. It is efficient, though the brevity borders on under-specification given the mutation risk.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation tool with no annotations, no output schema, and an undocumented parameter, the description is too thin. It omits the failure/edge behavior (unknown address, last recovery owner) that an agent would need before invoking it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single 'address' parameter has no schema description (0% coverage), so the description must compensate. It does add the meaning that the address is a 'saved recovery address' rather than an arbitrary string, but it does not clarify format, whether it must match an existing saved entry, or case sensitivity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Remove a saved recovery address'), so an agent knows exactly what operation it performs. It does not, however, distinguish itself from its obvious siblings set_recovery_owner and list_recovery_owners, which appear in the tool list and could be confused with it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus list_recovery_owners (to inspect addresses) or set_recovery_owner (to add one). There is no mention of prerequisites, such as whether the address must already be saved, or what happens if it is the last remaining recovery owner.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_auditBInspect
Run provenance audit on a Safe/address. Use when the user asks 'is this fake?' or wants a shareable verdict.
| Name | Required | Description | Default |
|---|---|---|---|
| chain | No | ||
| agents | No | Comma-separated agent keys | |
| address | Yes | Safe/address to audit (0x…40 hex) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full behavioral burden. It never states that this is a read-only/compute operation, whether anything is persisted or registered, what happens if the address is invalid or absent, or how costly/long the audit is. 'Audit' implies read semantics, but that is inference, not disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler; the operation is front-loaded and the usage cue follows immediately. Nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema means the description needn't describe the verdict payload, and the when-to-use cue is present. But with no annotations and partial schema coverage, it leaves gaps around read/write semantics, the meaning of 'agents', and chain defaults that an agent would want.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 67%, and the description adds nothing about any parameter. 'Safe/address' loosely echoes the address param, but chain, the enum values, and 'comma-separated agent keys' are left entirely to the schema. With a moderate coverage gap and zero compensating detail, this is below baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specifies a concrete verb and resource: run a provenance audit on a Safe/address. It reads clearly apart from the get_* siblings, though it never names an alternative tool or distinguishes itself from passport/alert lookups explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a real trigger condition ('is this fake?') and the intent behind it (shareable verdict), which is more than most siblings offer. It lacks any exclusion or routing to related tools, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_agent_passportAInspect
Declare who a watched agent is: model provider, run cadence, role, purpose, operator, and whether to publish a shareable public passport at /a/. Declarations never change the receipt-derived numbers next to them; publishing exposes only this address, never the rest of the watch list.
| Name | Required | Description | Default |
|---|---|---|---|
| role | No | ||
| model | No | Optional model or framework detail | |
| public | No | Publish a public passport page | |
| address | Yes | Watched agent address (0x…40 hex) | |
| cadence | No | always_on, scheduled, event_driven, on_demand | |
| purpose | No | One or two sentences on what it is for | |
| operator | No | Who runs it | |
| provider | No | Which model drives the agent | |
| display_name | No | Name shown instead of the raw address |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that declarations never alter receipt-derived numbers (a metadata-only write) and that publishing exposes only the single address and not the rest of the watch list. It omits auth/permission requirements and whether declarations are reversible or updatable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the purpose front-loaded before the behavioral caveats; every clause earns its place. The second sentence is dense but carries two distinct, useful guarantees.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter mutation with 89% schema coverage, no output schema, and no annotations, the description covers purpose, the metadata-only guarantee, and the public-publishing privacy scope. It leaves open permission needs and the return shape, but is otherwise adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 89%, so the schema already documents nearly every parameter, including enum values and descriptions. The description restates the field list at a high level without adding syntax or format meaning beyond the schema, matching the baseline for high-coverage schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ("Declare") and resource (who a watched agent is), then enumerates the exact attributes set (provider, cadence, role, purpose, operator, publishing). An agent can distinguish it from get_agent_passport by the verb, but the sibling add_watch that must precede it is never named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase "who a watched agent is" implies the agent must already be on the watch list, which is genuine pre-condition context. However, it never states when to call this versus add_watch or when to re-declare, so usage is inferred rather than specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_recovery_ownerAInspect
Save or update a recovery address for this user. Product default Safe is then 1/2: threshold 1, owners = [agent, this recovery]. Use make_default=true to prefer this label in get_safe_defaults.
| Name | Required | Description | Default |
|---|---|---|---|
| label | No | Short name, e.g. hardware, treasury, personal | |
| address | Yes | 0x recovery wallet the human controls | |
| purpose | No | What this recovery is for, e.g. default, trading | |
| make_default | No | Prefer this label for default Safes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations at all, the description carries the full burden, and it does disclose a non-obvious consequence: the product default Safe becomes 1/2 with threshold 1 and owners [agent, this recovery]. What it omits is what happens to a pre-existing recovery label with the same name (overwrite vs. add) and any permission/auth requirements for a write operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences, front-loaded with the core action and followed by the side effect and the flag semantics. Each sentence carries information; only the terse '1/2: threshold 1, owners = [...]' shorthand requires parsing effort.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter mutation tool with no annotations and no output schema, the description covers the action, the resulting Safe configuration, and the make_default behavior. It leaves open the overwrite/idempotency semantics and the return value, which are the remaining unknowns an agent might want.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented in the schema; the baseline is 3. The description only adds marginal meaning by connecting make_default to get_safe_defaults, which the schema's own text ('Prefer this label for default Safes') essentially covers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb pair and resource: 'Save or update a recovery address for this user,' which is clearly distinct from the sibling tools list_recovery_owners and remove_recovery_owner. It also states the downstream effect on the Safe, so an agent knows exactly what operation this is without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a concrete condition for the make_default flag ('Use make_default=true to prefer this label in get_safe_defaults'), which is real usage guidance tied to a named sibling. It does not, however, state when to prefer this tool over remove_recovery_owner or when re-calling it is inappropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
watch_demo_setAInspect
ONE-SHOT cold-start for brand-new Claude / MCP users with an empty watch list. Adds the three July 18 fixture addresses (Agent EOA + Safe A + Safe B on Base). Call this FIRST after connect when list_watched is empty, then call get_dashboard (Always allow the App). Idempotent — skips addresses already watched.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and does disclose key behavior: the operation is idempotent and skips addresses already watched, and it mutates the watch list by adding three specific entries. Permission behavior is only hinted at via the '(Always allow the App)' aside, and it doesn't say what the call returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the scenario ('ONE-SHOT cold-start for brand-new users with an empty watch list') and tight overall, with the address list and idempotency note each earning their place. The parenthetical UI instruction is slightly extraneous but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, parameterless mutation with no output schema, the description covers what it does, when to call it, the follow-up call, and idempotency. The main omission is any statement of side effects beyond the watch list or error/permission conditions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a parameterless tool applies. No parameter claims are made or needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — a one-shot seeding of the watch list with three named fixture addresses (Agent EOA, Safe A, Safe B on Base). It is clearly distinguishable from siblings like add_watch and list_watched, which handle individual additions and enumeration respectively.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit precondition is given: call this first, after connect, when list_watched is empty, and then call get_dashboard. The sequencing is unambiguous. It falls short of a 5 only because the alternative path (add_watch for non-empty or arbitrary lists) is implied rather than named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
21 tool updates
- First observed
add_watch - First observed
get_agent_passport - First observed
get_alerts - First observed
get_benchmark - First observed
get_capabilities - First observed
get_dashboard - First observed
get_feed - First observed
get_fidelity - First observed
get_leaderboard - First observed
get_pending_proposals - First observed
get_safe_defaults - First observed
get_tca - First observed
list_recovery_owners - First observed
list_watched - First observed
refresh_watches - First observed
register_intent - First observed
remove_recovery_owner - First observed
run_audit - First observed
set_agent_passport - First observed
set_recovery_owner - First observed
watch_demo_set
Related MCP Connectors
Read-only smart-contract security intelligence for autonomous agents.
Watchdog for unattended AI agents: alerts, evidence checks and a verifiable proof per run.
Pre-execution governance for AI agents. Deterministic PASS/FAIL/REVIEW verdicts, replayable proof.
DeFi safety layer for AI agents: wallet safety, token risk, tx decode/simulate. 20 tools.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceProvides cryptographic governance receipts for AI agents, enabling pre-execution evaluation and signed verdicts (EXECUTE/BLOCK/REVIEW/SHADOW) with offline-verifiable audit trails.MIT
- AlicenseAqualityCmaintenanceCryptographic accountability for AI agents. Ed25519-signed receipts for every MCP tool call. Constraints, chains, AI judgment, invoicing, and local dashboard included.2413 npm1MIT
- AlicenseNot gradedqualityDmaintenanceSecurity layer for AI agents that evaluates transaction intents and returns verdicts (ALLOW/WARN/DENY) using deterministic rules, on-chain checks, and simulation.9 npmMIT

evermint-mcpofficial
AlicenseNot gradedqualityDmaintenanceTamper-evident receipts for AI agent actions. The notary layer for agent-to-agent transactions.35 npm1MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.