proving-ground
Server Details
Accounts-payable exception environment with a pass/fail verifier and proof packet: AI agents work AP
Glama couldn't complete the latest health check. If this server requires authentication, missing or expired test credentials may be the cause. A test profile lets Glama authenticate for health checks and discover tools; it is separate from your personal connections.
If you are the author, claim ownership, then add or update a test profile under Admin → Test Profile.
- Status
- Unhealthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-06-18
- URL
- Repository
- bapatakshay1/proving-ground
- GitHub Stars
- 0
TDQS
Scored across 24 tools
Most tools target a clearly distinct resource+action (get_invoice vs get_po vs get_vendor, flag_vendor vs flag_bank_transaction, approve vs reject vs dispute vs hold). The many retrieval tools (get_*, list_open_exceptions, search_*) are differentiated by description, though the terminal-state invoice actions plus the resolve_exception wrapper require reading descriptions carefully to avoid misselection.
Strong, predictable verb_noun snake_case pattern (approve_invoice, search_pos, match_bank_transaction, list_open_exceptions). Minor deviations are the verb-only 'calculate' and the noun-only 'proof', which stand out but remain readable and unambiguous.
24 tools is on the heavy side, but the server spans several entities (invoices, POs, vendors, bank transactions, exceptions, policy, seat, proof), so most tools earn their place. It is slightly over the comfortable range but not bloated.
The AP exception lifecycle is well covered: retrieval, search, invoice state changes, exception resolution/escalation, reconciliation, and policy/proof support. Minor gaps exist—no bank-transaction search (only get/flag/match) and no vendor record update despite flag_vendor—but agents can work around these.
Available Tools
24 toolsapprove_invoiceApprove invoiceBDestructiveIdempotentInspect
Approve an invoice. approved_amount defaults to the invoice total; set it lower to short-pay.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| invoice_id | Yes | ||
| approved_amount | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (destructiveHint=true, idempotentHint=true), so the bar is lower. The description adds genuinely useful behavioral context — that approved_amount silently defaults to the invoice total and that a lower value constitutes a short-pay — which is not derivable from structured fields, though it omits any warning about the destructive/irreversible nature of the action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly written sentences with zero padding, and the core action is front-loaded ahead of the parameter detail. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with destructiveHint=true and no output schema, the definition covers the essential parameter behavior but omits the note parameter, any permission or state prerequisites, and any indication of what approval triggers downstream. Given the 0% schema coverage, it is adequate but leaves clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden. It competently explains the most ambiguous parameter (approved_amount's default and short-pay semantics) but says nothing about note, leaving one of three parameters completely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Approve an invoice"), which is unambiguous on its own. However, it never differentiates itself from the very similar siblings reject_invoice, hold_invoice, and dispute_invoice, so the agent must infer which action applies from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no alternatives named. The description never says when approval is appropriate versus rejection, holding, or disputing an invoice, nor what prerequisites (e.g. matched PO, no open exception) must exist first.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
calculateCalculate exactlyARead-onlyIdempotentInspect
Evaluate an arithmetic expression exactly (+ - * / parentheses, min, max, round, abs, sum of a list). Use it for every amount you derive: variances, min(invoiced, received) x price, tax, pair sums. Never do arithmetic in your head.
| Name | Required | Description | Default |
|---|---|---|---|
| expression | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this a pure, read-only, idempotent, non-destructive operation, so the safety burden is met. The description adds the valuable trait that evaluation is exact (avoiding float drift) and that no mental math should be trusted, but does not say how a malformed expression is handled.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: capability, grammar, then the usage imperative. The directive to use it is front-loaded after the capability, and no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, no-output-schema tool this is nearly complete: purpose, syntax, and usage pressure are all present. The only omission is what is returned (a value/error) and a concrete expression example, which would remove the last ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry the parameter meaning, and it does: it specifies the grammar (+ - * / parentheses, min, max, round, abs, sum of a list). It stops short of showing exact call syntax, e.g. how a list is delimited for sum, leaving a small inference gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (evaluate) and resource (an arithmetic expression) and enumerates supported operators/functions, so the agent knows exactly what this tool computes. It is unmistakably distinct from every sibling, which are all invoice/vendor/bank workflow tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit directive with conditions: 'Use it for every amount you derive' and 'Never do arithmetic in your head,' plus concrete example cases (variances, min(invoiced, received) x price, tax, pair sums). No alternative could be confused here, so the absence of a named sibling alternative is not a gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dispute_invoiceDispute invoiceCDestructiveIdempotentInspect
Dispute an invoice with the vendor.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| reason | Yes | ||
| invoice_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, openWorldHint=false, and readOnlyHint=false, so the safety profile is partly covered. The description adds nothing on top: it never says what disputing does (vendor notified? invoice locked? state transition?), whether it is reversible, or what permissions are needed for an irreversible destructive action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is a single short sentence with no filler, but that brevity here is under-specification rather than economy. The one clause that exists is front-loaded correctly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, state-changing invoice operation with zero parameter documentation and no output schema, the description is materially incomplete. An agent cannot tell what side effects to expect, which makes invoking an irreversible action risky.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden and fails it. Neither the required 'reason' nor the optional 'note' is explained (allowed values, length, how the reason is conveyed to the vendor), and 'invoice_id' is only implied by the phrase 'an invoice'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description has a specific verb and resource ('Dispute an invoice'), and 'with the vendor' hints at the external party involved. It does not, however, distinguish this action from the sibling lifecycle tools (approve_invoice, reject_invoice, hold_invoice), so an agent must infer the difference from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no preconditions, and no mention of alternatives such as approve_invoice, reject_invoice, or hold_invoice that also act on invoices. The agent gets no rule for choosing dispute over rejection or escalation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
escalate_exceptionEscalate exception (metered claim)BDestructiveInspect
Hand the exception to a human because the policy cannot be applied.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| exception_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructiveHint=true, idempotentHint=false and readOnlyHint=false, so the safety profile is covered structurally. The description adds the useful semantic outcome that the case leaves automated handling and goes to a human, but it does not say what state the exception lands in, whether repeating the call creates duplicate escalations, or whether it can be undone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the action, the recipient, and the triggering condition with zero filler. Nothing could be trimmed without losing the rationale clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent, two-parameter mutation with no output schema and 0% parameter coverage, the description is too thin. It omits parameter meaning and the observable result of escalation, so an agent cannot confirm what happens to the exception after the call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so both exception_id and reason are undocumented in the schema, and the description supplies nothing about their format or content. It never hints at what a valid exception_id looks like or what a useful reason string should contain, leaving the burden entirely unmet.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and object: routing an exception to a human. It implicitly differentiates from resolve_exception by framing this as the path taken when policy cannot be applied, but it never names the sibling it displaces, so an agent must infer the split.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"because the policy cannot be applied" gives an implied precondition for use, which is more than nothing, but there is no explicit when-not guidance and no named alternative such as resolve_exception or reject_invoice for cases where policy does apply.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flag_bank_transactionFlag bank transactionCDestructiveIdempotentInspect
Flag a bank transaction for treasury review.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| bank_transaction_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true and readOnlyHint=false, so the agent knows this mutates state, but the description adds nothing about what flagging changes, whether it can be undone, or who sees the flag. For a destructive write with annotation coverage already thin on specifics, the description should say more than 'flag for review'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, which is structurally sound. However, its brevity is under-specification rather than discipline — it omits information the agent needs, so it is not efficient in the sense of covering the task.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a destructive, non-read-only mutation with two undocumented parameters and no output schema. The description omits the effect of the flag, reversibility, and any permission or workflow context, leaving the agent to guess at the tool's consequences.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not compensate: it never explains what the 'reason' should contain, whether it is free text or a coded value, or what the transaction ID format is. The parameter names are self-explanatory enough to avoid a 1, but no added meaning is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Flag') and resource ('bank transaction') plus the downstream purpose ('for treasury review'), which separates it from get_bank_transaction and match_bank_transaction. It stops short of naming any sibling or contrasting behavior explicitly, so it is clear but not differentiating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to flag versus escalating (escalate_exception) or disputing, no prerequisites, and no indication of what happens after flagging. The 'for treasury review' phrase implies a workflow but never states the trigger conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flag_vendorFlag vendorCDestructiveIdempotentInspect
Flag a vendor master record for verification.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| vendor_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered. The description adds nothing about what flagging actually does to the vendor record, whether it blocks payments or triggers downstream review, or how to un-flag it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the action and target front-loaded and no filler. It is efficient, though the brevity comes at the cost of the missing detail noted elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-read-only mutation with no output schema and zero parameter documentation, the description is too thin. It omits the effect of the flag, reversibility, and the semantics of the required 'reason' argument.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for two undocumented parameters, but it says nothing about vendor_id format or what the 'reason' string should contain. Baseline of 3 is not met because the description adds no parameter meaning beyond the bare names in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Flag') and resource ('vendor master record') plus the intent ('for verification'), which cleanly separates it from read-oriented siblings like get_vendor and search_vendors. It does not, however, name or differentiate itself against other mutating siblings such as escalate_exception or reject_invoice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to flag a vendor versus escalating an exception or disputing an invoice, and no stated prerequisites or trigger conditions. The agent must infer the workflow context entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_bank_transactionGet a bank transactionCRead-onlyIdempotentInspect
A bank transaction.
| Name | Required | Description | Default |
|---|---|---|---|
| bank_transaction_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered structurally. The description adds no behavioral context of its own — no return shape, no lookup-failure behavior, no indication of what the id resolves against — so it contributes essentially nothing beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The text is short, but this is under-specification rather than conciseness — a six-word fragment that conveys no actionable information. Brevity only earns credit when the content is complete.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-parameter lookup with no output schema, the description should at minimum state what is retrieved and how the id is used. Nothing about the return value or failure modes is provided, leaving the definition materially incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions `bank_transaction_id` or what kind of transaction identifier is expected. The single parameter's meaning is only inferable from its self-evident name, so the description fails to compensate for the documentation gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"A bank transaction." is a noun phrase that simply restates the title and tool name without stating a verb or scope. An agent learns nothing it could not infer from `get_bank_transaction` alone, and there is no differentiation from siblings like `match_bank_transaction` or `flag_bank_transaction`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as `match_bank_transaction` or `flag_bank_transaction`. The agent must infer entirely from the tool name whether retrieval, matching, or flagging is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_exceptionGet an exceptionDRead-onlyIdempotentInspect
An exception from the queue.
| Name | Required | Description | Default |
|---|---|---|---|
| exception_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, so the safety profile is covered. The description adds only the vague phrase 'from the queue' and says nothing about lookup failures, permissions, or what a missing exception returns, so it contributes almost no context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Brevity here reflects under-specification rather than conciseness: it is a fragment with no verb and nothing front-loaded of value. One complete sentence stating what is retrieved would cost no more space and be far more useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-by-ID tool the annotations cover the safe-read profile, but with no output schema the description should at least characterize what an exception record is and what is returned. It leaves the agent with only the tool name to work from.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single required exception_id parameter, so the description carries the burden of explaining it and fails to — no format, no source of the ID, no relationship to the queue. It is marginally better than the zero-coverage multi-param case only because a single 'get by ID' parameter is partly self-evident.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a noun fragment, 'An exception from the queue,' that essentially restates the title 'Get an exception' without a verb or scope. It does not distinguish this tool from siblings like list_open_exceptions, resolve_exception, or escalate_exception, so an agent cannot tell which retrieval semantics apply.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no mention of alternatives such as list_open_exceptions for enumeration or resolve_exception/escalate_exception for acting on an exception. The agent is given nothing to route on beyond the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_invoiceGet an invoiceCRead-onlyIdempotentInspect
An invoice with its lines, vendor, and linked PO (PO lines include qty_received and qty_invoiced).
| Name | Required | Description | Default |
|---|---|---|---|
| invoice_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so the safety profile is covered. Because there is no output schema, the description's disclosure of return contents (lines, vendor, linked PO, and PO-line fields qty_received/qty_invoiced) adds real value, though it says nothing about not-found behavior or payload size.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no padding, but it is an incomplete fragment lacking a subject-verb clause, so brevity here reflects under-specification rather than tight editing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema the description does useful work by sketching the return shape, but it omits the lookup-by-ID framing, the invoice_id parameter, and error behavior — gaps that matter for a single-record fetch tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
One required parameter, invoice_id, with 0% schema description coverage, and the description never mentions it. The agent must infer from the parameter name alone that this is the invoice identifier and that it is mandatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a noun phrase ('An invoice with its lines, vendor, and linked PO') with no verb, so the retrieval action is only implied by the name get_invoice. It conveys what the returned object contains but never states that the tool fetches a single invoice by ID, and it draws no distinction from siblings like search_invoices or get_po.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus search_invoices (list/filter) or get_po (linked purchase order). Nothing tells the agent the invoice_id is the lookup key or what happens if the ID is unknown.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_poGet a purchase orderCRead-onlyIdempotentInspect
A purchase order with lines (qty_received, qty_invoiced) and linked invoice ids.
| Name | Required | Description | Default |
|---|---|---|---|
| po_number | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive, closed-world behavior, so the safety profile is covered. The description adds the return shape (lines with qty_received/qty_invoiced plus linked invoice ids), which is useful given there is no output schema, but it says nothing about failure behavior when the PO doesn't exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the most important content (what the PO object contains). It wastes no words, though the fragmentary phrasing means little is actually conveyed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read whose annotations cover the safety profile, the description partially compensates for the absent output schema by outlining the returned PO structure. However, it omits the parameter's meaning and any failure/not-found behavior, leaving meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single parameter, po_number, and the description never mentions it or clarifies its format or source. The parameter is self-evidently named, which limits the damage, but the description fails to compensate for the documentation gap as required at this coverage level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is a noun-phrase fragment ('A purchase order with lines...') that describes the resource's contents rather than stating a verb like 'retrieves a single purchase order by number.' The title supplies the verb, but the description itself doesn't clearly distinguish this single-PO fetch from siblings such as search_pos or link_po.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of alternatives like search_pos, and no note about prerequisites such as needing a valid po_number. The agent is left to infer the tool's role from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_policyGet the AP exception policyARead-onlyIdempotentInspect
The AP exception policy (AP-POL-7). Read it before deciding.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint=false, so the safety profile is fully covered. The description adds modest context by identifying the document as a specific policy artifact (AP-POL-7), but discloses nothing further about content, caching, or freshness.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler, and the identification of the resource is front-loaded. It is efficient, though the second sentence is terse enough to be slightly cryptic.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-param, read-only document getter with annotations covering safety, the definition is nearly sufficient. The one gap is that, absent an output schema, it does not say what is returned (e.g., policy text vs. structured rules), leaving the agent to discover the shape on first call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing to disambiguate; per the baseline for a 0-param tool this scores 4. No parameter explanation is needed or expected.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and a specific resource (the AP exception policy), even supplying a document identifier (AP-POL-7). It is clearly distinguishable from sibling lookups like get_exception or get_invoice, though it does not explicitly name those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Read it before deciding' implies the tool should be consulted before adjudicating exceptions, which is genuinely useful guidance. However, it never names which decisions or sibling tools (approve_invoice, reject_invoice, escalate_exception) it should precede, so the routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_seatMy seat and creditsBRead-onlyIdempotentInspect
Who you are, your sandbox, credits left and price per claim.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered without the description. The description adds value by enumerating the returned information (identity, sandbox, credits left, price per claim), which compensates for the absence of an output schema. It does not, however, discuss caching, rate limits, or when credits are refreshed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with no filler, front-loaded with the most important payload (who you are). It is terse to the point of being a telegraphed field list rather than prose, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no input parameters, no nested objects, and no output schema, the description carries the burden of indicating what comes back; listing identity, sandbox, credits left, and price per claim largely does that. It is adequate for the tool's simplicity, though an agent still cannot know the exact response shape or units for 'price per claim'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4 and there is nothing to disambiguate. The description correctly implies a parameterless, context-derived call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific resource it returns – the caller's identity, sandbox, remaining credits, and per-claim price – which maps cleanly onto the title 'My seat and credits'. It is distinguishable from every sibling, none of which deal with account/seat state. The verb 'get' is implied rather than stated, and the sentence is a noun list rather than a verb+resource phrase.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no mention of prerequisites or when this should be preferred over siblings such as get_policy or search_invoices. Usage is only inferable from the name (fetch my own account state). For a zero-argument tool this is a minor gap, but nothing in the text directs the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_vendorGet a vendorCRead-onlyIdempotentInspect
Vendor master record.
| Name | Required | Description | Default |
|---|---|---|---|
| vendor_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=false, so the safety profile is fully covered structurally. The description adds nothing beyond them – no note on return shape, missing-vendor behavior, or permission requirements – so it contributes almost no behavioral value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three words with no waste, but this is under-specification rather than conciseness – the brevity removes information an agent needs rather than trimming filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even for a simple single-parameter read with no output schema, the description omits what is returned and what identifies a valid vendor, leaving the agent dependent entirely on annotations and the bare parameter name.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the sole parameter vendor_id carries no description anywhere. The phrase "vendor master record" implies a lookup key but gives no format, source, or semantics for the id, so the description fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Vendor master record" is a noun phrase that essentially restates the name/title (get_vendor) without stating a verb or scope. It hints at the resource but never says it retrieves a single vendor by identifier, nor distinguishes it from siblings like search_vendors.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance whatsoever. The description does not mention search_vendors as the alternative for discovery-style lookups, nor flag_vendor for state changes, leaving the agent to infer routing from names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hold_invoiceHold invoiceCDestructiveIdempotentInspect
Put an invoice on hold with a reason code.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| reason | Yes | ||
| invoice_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered structurally. The description adds nothing beyond that: it does not say whether the hold is reversible, what state the invoice enters, whether the note is stored, or what happens on repeat calls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no waste, but its brevity reflects under-specification rather than disciplined editing — it omits information the agent needs for a destructive mutation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-read-only state-changing tool with 0% schema description coverage and no output schema, the description is too thin: no effect description, no permission or prerequisite info, no guidance on the note field, and a misleading 'reason code' framing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry parameter meaning, and it does not. It mentions 'reason code' (suggesting a fixed enumeration) while the schema defines reason as an unconstrained string with no enum, which is actively misleading; invoice_id and note are never addressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('put on hold') and resource ('invoice'), which is distinct from approve_invoice and reject_invoice by name. However, it offers no explicit differentiation from siblings like dispute_invoice or escalate_exception, which are plausible alternative state transitions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to hold versus dispute, reject, or escalate an invoice, and no prerequisites or conditions. The agent must infer usage entirely from the tool name and sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
link_poLink invoice to POCDestructiveIdempotentInspect
Link an invoice to a purchase order number.
| Name | Required | Description | Default |
|---|---|---|---|
| po_number | Yes | ||
| invoice_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the agent knows this mutates state and is repeatable. The description adds nothing beyond that — it doesn't say what is destroyed or overwritten, whether an existing PO link is replaced, or whether the action can be undone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is appropriately sized, though its brevity reflects under-specification rather than disciplined editing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-readonly mutation with 0% parameter documentation and no output schema, the description is too thin: it omits what linking changes, side effects, and any preconditions the agent should check.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for both parameters, so the description must compensate. It only loosely maps the two args ('invoice' -> invoice_id, 'purchase order number' -> po_number) and gives no format, ID-space, or validation detail for either.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Link') and both resources ('invoice', 'purchase order number'), which is clear enough to distinguish from siblings like approve_invoice or match_bank_transaction. It does not, however, contrast itself with any sibling or clarify what 'link' means operationally.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites (e.g., invoice must exist or be unlinked), and no mention of alternatives despite a dense sibling set. The agent must infer everything from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_open_exceptionsList open exceptionsARead-onlyIdempotentInspect
Open exceptions in your sandbox, optionally filtered by kind (price_mismatch, quantity_mismatch, possible_duplicate, missing_po, unmatched_payment, vendor_bank_change). Start here.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered. The description adds value beyond them by disclosing the sandbox scope and enumerating the six exception kinds, which is behaviorally important for interpreting results. The undocumented 'limit' and lack of pagination/result-size behavior keep it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the scope and entry-point cue front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema and safety already covered by annotations, the description supplies the essential missing pieces: the kind vocabulary and the sandbox scope. Only the 'limit' parameter's semantics are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load. It compensates well for 'kind' by listing all six allowed values (the schema declares no enum), but leaves 'limit' completely unexplained, so half the parameters remain undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('Open exceptions') with scope ('in your sandbox') and a modifier ('optionally filtered by kind'). It is clearly distinguishable from siblings like list_resolved_examples and get_exception, which fetch a different state or a single item.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Start here.' explicitly signals this is the entry-point tool for the exception workflow, which is meaningful routing guidance. It stops short of naming when-not-to-use or pointing to a follow-up tool (e.g. resolve_exception), so it is clear context without exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_resolved_examplesList resolved examplesBRead-onlyIdempotentInspect
Recently resolved exceptions of a kind, with the clerk's resolution summary.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | ||
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive, and closed-world, so the safety profile is covered. The description usefully adds that results include 'the clerk's resolution summary' (valuable since there is no output schema), but 'recently' is an undefined window and pagination/ordering behavior is unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the resource and filter come first. It is efficient, though arguably too terse for the gaps it leaves open.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read tool whose annotations cover safety, the description is minimally adequate, but with no output schema and 0% schema coverage it should at least hint at what 'kind' accepts and what 'recently' means.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the parameter burden. 'of a kind' loosely maps to the required kind parameter but never explains valid kinds (no enum exists), and the limit parameter is not mentioned at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The phrase 'Recently resolved exceptions of a kind' gives a specific verb (list), resource (resolved exceptions), and scope (recency + kind filter), which implicitly contrasts with the sibling list_open_exceptions. It stops short of naming that sibling explicitly, so differentiation is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description is a noun phrase, not guidance: it never states when to reach for this tool versus get_exception, list_open_exceptions, or resolve_exception. There are no prerequisites, no exclusions, and no alternative named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
match_bank_transactionMatch payment to invoicesBDestructiveInspect
Reconcile a bank transaction against one or more approved invoices (they become paid).
| Name | Required | Description | Default |
|---|---|---|---|
| invoice_ids | Yes | ||
| bank_transaction_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and openWorldHint=false, so the agent knows this mutates state and is not repeatable. The description usefully adds the side effect that matched invoices become paid, but omits whether the match is reversible, how partial amounts are handled, and what authorization is required.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence front-loads the action and resource with no filler. It is perhaps overly terse for a destructive operation, but every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent mutation with no output schema and zero schema coverage, the description conveys the core action and the key outcome (invoices become paid). It still leaves gaps around reversibility, partial matching, and error conditions that an agent would want before invoking.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the two required parameters rely entirely on the description. The prose maps naturally to both (a bank transaction and one or more approved invoices) and adds the constraint that invoices must be approved, but it gives no format or cardinality detail beyond 'one or more'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (reconcile/match) and resources (bank transaction, invoices) and clarifies the outcome ('they become paid'). It distinguishes itself reasonably from read-side siblings like get_bank_transaction and from invoice-specific actions like approve_invoice, though it never names an alternative directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to choose this over siblings such as link_po, resolve_exception, or flag_bank_transaction, nor any stated preconditions beyond 'approved invoices'. The reconciliation context is implied but no exclusion or alternative routing is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
proofMy proof packetARead-onlyIdempotentInspect
Your proof packet: pass rate per workflow on the cases you claimed, with 95% CI and the human unit cost.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds the return shape (pass rate with CI, cost figure), which matters because there is no output schema, but it says nothing about data freshness, scope, or whose cases are included. Adequate but thin against the annotated baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence, front-loaded with the deliverable and followed by its exact contents. There is no filler, hedging, or repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and no output schema, the description must carry the return-value burden, and it does specify the three measurements returned plus their statistic (95% CI). What is missing is scope framing — what "claimed" cases are and over what period. Reasonably complete for a trivial-signature tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, which is the baseline-4 case. The description correctly adds no parameter detail and instead spends its words on payload semantics. Nothing misleading is introduced.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the concrete deliverable ("your proof packet") and enumerates its contents: pass rate per workflow, 95% CI, and human unit cost. That is far more specific than the bare name "proof" and tells an agent what it will get. It does not, however, distinguish this from any sibling such as list_resolved_examples or calculate, so a 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use statement, no prerequisites, and no named alternative. The phrase "the cases you claimed" hints at a post-hoc verification scenario, but the agent must infer when this tool is the right choice versus list_resolved_examples or calculate. Implied-only guidance merits a 2.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reject_invoiceReject invoiceDDestructiveIdempotentInspect
Reject an invoice.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| reason | Yes | ||
| invoice_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=true, and idempotentHint=true, so the safety profile is covered structurally. The description adds nothing beyond that: it does not say the rejection is terminal, whether it can be undone, who is notified, or what the reason field should contain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single four-word sentence is short but not concise in any useful sense; it is under-specified rather than economical. Nothing is front-loaded because nothing is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a destructive, non-read-only mutation with three undocumented parameters, no output schema, and no annotation-independent context in the description. An agent cannot call it correctly on the description alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across three parameters, and the description supplies no meaning for invoice_id, reason, or note. The distinction between required reason and optional note is left entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Reject an invoice" restates the tool name and title almost verbatim, adding no scope, no distinguishing detail. It does not separate this from close siblings like approve_invoice, dispute_invoice, or hold_invoice, so the agent gets no basis for choosing among them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance whatsoever on when to reject versus dispute, hold, or escalate an invoice, nor any prerequisite or precondition. The agent must infer usage entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resolve_exceptionResolve exception (metered claim)BDestructiveInspect
Mark the exception resolved. Call this last, after the record is in its final state.
| Name | Required | Description | Default |
|---|---|---|---|
| summary | Yes | ||
| exception_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, so the agent knows this is an irreversible state change. The description's 'call this last, after the record is in its final state' adds useful sequencing context, but says nothing about what is destroyed, whether re-resolution is possible, or what the result looks like.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, no filler, with the action stated first and the ordering constraint second. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, non-idempotent mutation with two required yet completely undocumented parameters and no output schema, the description is too thin. An agent cannot tell what to put in 'summary' or confirm what 'resolved' entails beyond the annotation flags.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and both parameters are required, so the description carries the burden of explaining them — yet it never mentions exception_id or summary. In particular, what 'summary' should contain (a resolution note? a computed total?) is left entirely undefined.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Mark the exception resolved.' The action is distinguishable from sibling escalate_exception, though the description never names siblings or contrasts the two terminal-state operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Call this last, after the record is in its final state' gives a clear ordering precondition for when to invoke it. It stops short of naming alternatives (escalate_exception, dispute_invoice) or stating when not to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_invoicesSearch invoicesARead-onlyIdempotentInspect
Search invoices. All filters optional; omit the ones you do not need. Results are capped by limit (default 50, max 200), newest first, so raise the limit or add filters (invoice_number, min_total/max_total, date_to) before concluding something does not exist.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| status | No | ||
| date_to | No | ||
| date_from | No | ||
| max_total | No | ||
| min_total | No | ||
| po_number | No | ||
| vendor_id | No | ||
| invoice_number | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), and the description adds genuinely new behavioral detail beyond them: the default (50) and maximum (200) result cap, newest-first ordering, and the false-negative trap. That is exactly the context an agent needs and cannot get from the annotations or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loaded with the core action and then the essential caveat. Every clause earns its place; nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-output-schema search tool, it covers the decisive behavioral facts (cap, ordering, false-negative avoidance). It is slightly incomplete on the undocumented filter parameters and their accepted formats, but the practical guidance is strong.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load. It usefully documents limit (default 50, max 200) and names invoice_number, min_total/max_total, and date_to, but leaves status, date_from, po_number, and vendor_id — half the parameters — completely unexplained, including any accepted values or formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource ("Search invoices") that an agent can act on immediately. It does not explicitly differentiate from siblings like search_pos or search_vendors, though the resource name makes the scope self-evident.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives actionable guidance: all filters are optional, omit unneeded ones, and — critically — raise the limit or add filters before concluding something does not exist. This directly addresses a real failure mode. It stops short of naming alternative sibling tools for related lookups.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_posSearch purchase ordersARead-onlyIdempotentInspect
Purchase orders for a vendor, optionally filtered by status (open|closed) and by a SKU that appears on the PO lines. Capped by limit (default 50, max 200), newest first; filter by sku to find the PO for an invoice line.
| Name | Required | Description | Default |
|---|---|---|---|
| sku | No | ||
| limit | No | ||
| status | No | ||
| vendor_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds real behavioral value beyond that: a result cap (default 50, max 200) and deterministic ordering (newest first). It does not describe pagination or the returned PO shape, but the core behavioral traits are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with zero filler, front-loading the resource and scope before the filters and the invoice-line use case. Every clause conveys information the agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description needn't detail return fields, and it adequately covers filtering, ordering, and result limits. Minor gaps are the unpaginated cap behavior and the lack of explicit routing to sibling retrieval tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the full burden and largely meets it: it defines the status enum (open|closed), explains sku as "a SKU that appears on the PO lines," and gives limit's default and maximum. vendor_id is only implied by "for a vendor," so one parameter remains slightly under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and scope ("Purchase orders for a vendor") plus the two optional filters, so the agent immediately knows this returns a filtered set rather than a single PO. The vendor-scoped, multi-result framing implicitly separates it from the sibling get_po which fetches one PO.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a concrete when-to-use scenario ("filter by sku to find the PO for an invoice line"), which is genuinely helpful context. It stops short of naming alternative siblings (get_po, search_invoices) or stating when-not to use this tool, so it is clear context without explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_vendorsSearch vendorsARead-onlyIdempotentInspect
Find vendors by (partial, case-insensitive) name, e.g. a bank counterparty string.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=true, idempotent=true, destructive=false and closed-world, so the safety profile is covered. The description adds genuinely useful behavioral semantics — partial and case-insensitive matching — but says nothing about result caps, ordering, or behavior on zero matches.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with the matching semantics and a concrete example; no filler, nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter lookup with full annotation coverage and no output schema, the description supplies the one non-obvious detail (fuzzy name matching). Result-shape expectations are the only unaddressed item, and the absence of an output schema keeps that a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single parameter is an undocumented plain string, so the description carries the burden — and it does, specifying that the match is partial and case-insensitive, which an agent needs to know before invoking. It stops short of noting whether an empty string is valid or how multiple matches behave.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Find) plus resource (vendors) and even the matching mode (partial, case-insensitive), which is more than a bare restatement. It does not distinguish itself from the sibling get_vendor, which is the only missing piece for a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The example ("a bank counterparty string") implies the intended use case of resolving an unstructured name to a vendor, but there is no explicit guidance on when to use this versus get_vendor or the other search_* siblings. Usage must be inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
24 tool updates
- First observed
approve_invoice - First observed
calculate - First observed
dispute_invoice - First observed
escalate_exception - First observed
flag_bank_transaction - First observed
flag_vendor - First observed
get_bank_transaction - First observed
get_exception - First observed
get_invoice - First observed
get_po - First observed
get_policy - First observed
get_seat - First observed
get_vendor - First observed
hold_invoice - First observed
link_po - First observed
list_open_exceptions - First observed
list_resolved_examples - First observed
match_bank_transaction - First observed
proof - First observed
reject_invoice - First observed
resolve_exception - First observed
search_invoices - First observed
search_pos - First observed
search_vendors
Related MCP Connectors
Pre-payment checks for AI agents: x402 quotes, untrusted prompts and audit receipts.
Agent-native double-entry accounting ledger with x402 micropayments
Verifiable provenance for AI agents — ZK proofs over confidential documents, no plaintext exposure.
Verifier-grounded AI promotion gates, disposable report cards, and signed PASS/HOLD/BLOCK receipts.
Related MCP Servers
- AlicenseAqualityBmaintenanceEnables LLM agents to query and reconcile an accounts-payable ledger through safe read-only SQL, deterministic invoice-to-payment matching, and exception explanations for underpayments, duplicates, currency mismatches, and overdue invoices. Ships with a synthetic ledger, golden-set evals, and a month-end review workflow so quality can be measured in CI.4MIT
- FlicenseAqualityBmaintenanceEnables an AI agent to handle accounts payable tasks against a mock ERP, including reading and writing bills and vendors, checking duplicates, matching invoices, recommending approvals, and queuing payment releases, with configurable profiles that limit available tools.11-
- AlicenseNot gradedqualityBmaintenanceAn accounting-ops agent that reconciles payments against open orders, auto-books provably safe payments through a deterministic policy gate, and escalates exceptions to a human queue with audit trails.MIT
- AlicenseAqualityDmaintenanceAudit infrastructure for AI agents to log consequential decisions (invoice, GL, anomaly) and verify attestations via MCP tools.6MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.