tillbooks
OfficialTILL
Trusted Independent Ledger Library. Swiss accounting your agent can actually use.
TILL is a free, open-source (MIT), local-first accounting app built for Switzerland. Your books live in a single SQLite file on your own machine. They never leave the country, because they never leave your laptop.
The difference is the interface. Most accounting software is a GUI with an AI button bolted on. TILL is built the other way around: an MCP server is a first-class interface, so an agent can post entries, categorize transactions, draft and chase invoices, prepare the MWST-Abrechnung, and answer "how did my quarter go" by talking to the ledger directly. A minimalist Studio gives you human oversight of everything the agent did. Agent and human share one ledger.
Status: pre-alpha. The engine, the MCP verbs and the Studio are built across the numbered capabilities. Nothing here is ready to keep real books yet. Do not run your business on it.
Why
Switzerland has accounting software. It does not have accounting software an agent can drive, and the open-source option (Gäld) is AGPL, which forecloses an open-core model. TILL is a clean-room MIT rebuild aimed at the things that actually make Swiss books Swiss:
QR-bill (Swiss QR-Rechnung) with the structured-address standard
Three MWST rates (8.1%, 3.8%, 2.6%) and the MWST-Abrechnung that falls out of them
Kontenrahmen KMU seeded out of the box
ISO 20022 (camt/pain) for banking
de-CH and en from day one
Related MCP server: bookie
The ledger is not negotiable
Posted entries are append-only and immutable. A correction is a reversing entry, never a destructive edit, exactly as real accounting works. Posting is idempotent: a double-post does not double-count. These are asserted in the test suite, not merely promised in a README.
Install and run
TILL is on npm as tillbooks (0.1.0, pre-alpha). Installing it globally puts the till command
on your path:
npm install -g tillbooks
till helpTo connect an MCP client such as Claude Desktop or Cursor, point it at the stdio server. No API key is needed:
{
"mcpServers": {
"tillbooks": {
"command": "npx",
"args": ["-y", "tillbooks", "mcp"]
}
}
}The ledger is created at ~/.till/till.db; set TILL_DB_PATH to put it somewhere else. till up
starts the Studio, the human oversight GUI.
To work on TILL itself, build it from a checkout:
npm install
npm run build # tsc emits dist/
node bin/till.mjs helpCONTRIBUTING.md has the full dev setup, and the quickstart walks the same steps.
Documentation
Quickstart : where TILL stands, and how it connects to your agent
Run inside an agent : drive TILL over MCP with no machine of your own
Self-hosting in under an hour : stand up a served instance
Documentation : the product guides and capability reference
Design : the design law this project builds to
Before you use it
TILL is pre-alpha and it keeps books. DISCLAIMER.md is worth two minutes: TILL is not tax advice, you stay responsible for your own filings, and you should not point it at real accounts yet.
Contributing
CONTRIBUTING.md : how to run it, and the rules that are not negotiable
CODE_OF_CONDUCT.md : how we treat each other
SECURITY.md : report vulnerabilities privately to security@tillbooks.ch
SUPPORT.md : where to ask questions
License
MIT, Copyright (c) 2026 Nomadik GmbH. See LICENSE.
TILL is an independent project. It is not affiliated with, endorsed by, or derived from bexio AG or any other accounting vendor. It is a clean-room implementation built from public documentation.
Available Tools
779 toolsaccept_inviteA
Redeem an invite token: bind this session actor to the invited identity and activate the membership. Pre-workspace, because the accepter is not a member yet.
| Name | Required | Description | Default |
|---|---|---|---|
| token | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It does disclose a key behavioral consequence: binding the session actor to the invited identity and activating membership, which implies a state change and identity mutation. However, it doesn't mention whether this is irreversible, whether the token is consumed, what errors occur on invalid/expired tokens, or whether this affects the current session permanently. Some gaps remain but the core side effect is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The core action and effect are front-loaded, and the contextual note about pre-workspace is succinctly placed. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one required parameter and no output schema, the description is nearly complete. It explains what happens, when it happens (pre-workspace), and the binding effect. It could be more complete on error cases or token lifecycle, but those are secondary for basic invocation. The context note about the accepter not being a member yet is useful and somewhat unusual. It's slightly incomplete on behavioral edge cases, but overall strong for this complexity level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only mentions "token" implicitly via "invite token." It doesn't specify the token's format, source, or that it's the single required parameter. The term "invite token" does echo the parameter name and gives limited semantic meaning, but an agent might benefit from knowing where to obtain the token or any length/encoding constraints. This earns a minimal 3 rather than lower because the description does tie the parameter to the redeem action.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (Redeem), the resource (invite token), and the effect (bind this session actor to the invited identity and activate the membership). The additional note "Pre-workspace, because the accepter is not a member yet" gives important context that disambiguates it from other invite/member tools like invite_member or revoke_member.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the context for using this tool (pre-workspace, accepter not yet a member), which helps an agent decide when to invoke it. However, it does not explicitly mention when not to use it or name alternative tools, so it falls short of full exclusion guidance. It's clear enough for an agent to identify the right scenario but lacks explicit alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
account_set_tax_defaultC
Set (or clear) an account's default tax code.
| Name | Required | Description | Default |
|---|---|---|---|
| taxCode | No | ||
| accountId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the full behavioral disclosure burden. It does reveal that calling the tool can either set or clear the default, but it omits details such as whether clearing is achieved by omitting the optional taxCode, whether the change is reversible. For a mutation tool with zero annotation coverage, this is a meaningful behavioral gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one efficient sentence with no filler and the core operation is front-loaded. It is concise, though it sacrifices important behavioral and parameter details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given three parameters, no output schema, and no annotations, the description is not complete enough for an agent to call correctly. The main gap is the mechanism for clearing the default tax code: the schema marks taxCode optional, but the description never says that omitting taxCode clears the value. It is an adequate statement of purpose, not a complete operational definition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds only a general reference to the default tax code. It does not explain the taxCode parameter's role beyond implying it is the value to set, and it does not say how clearing is indicated given taxCode is optional. The schema properties are named clearly, but at 0% coverage the description does not compensate enough for undeclared semantics such as how clearing works.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Set or clear) with a specific resource (an account's default tax code. This precisely states what the tool does and differentiates it from related but distinct tools such as update_account and set_vat_method. The scope is clear enough that an agent selecting among the tax and account siblings can identify it without inspecting the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance about how to choose this tool over alternatives, and it does not mention alternatives such as set_vat_method or update_account. There is also no indication of when one should clear versus set a tax default rather than use another VAT configuration tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
accrual_createA
Describe an Abgrenzung (OR Art. 958b) and save it as a DRAFT: kind is one of prepaid_expense (Aufwand vorausbezahlt, Dr 1300), accrued_income (Ertrag noch nicht fakturiert, Dr 1300), accrued_expense (Aufwand noch nicht fakturiert, Cr 2300) or deferred_income (Ertrag vorausbezahlt erhalten, Cr 2300); contraAccount is the P&L account (number or id, income for the income kinds, expense for the expense kinds); amountMinor is integer Rappen; periodEnd is the balance-sheet date. Returns the draft with the exact lines the post will write and the reversal lines dated the day after. Nothing is posted.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | ||
| currency | No | ||
| periodEnd | Yes | ||
| sourceRef | No | ||
| amountMinor | Yes | ||
| description | Yes | ||
| workspaceId | Yes | ||
| costCenterId | No | ||
| contraAccount | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses key behaviors: saves as a draft, returns the exact lines the post would write, includes reversal lines dated the day after, and explicitly states nothing is posted. This covers the most critical side effects, though it does not address idempotency or error handling, which are less critical for a draft action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence that front-loads the purpose and then explains parameters in a list-like fashion. Every phrase adds value, but the density might make it slightly harder to parse. It is concise without fluff, though structuring as bullet points could improve readability while retaining brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 10 parameters, no output schema, and no annotations, the description covers the essential aspects: the main behavior, parameter semantics for the most critical fields, and the expected return (draft lines and reversal lines). It does not explain all parameters, but the missing ones are peripheral. The explicit note that nothing is posted adds important context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It provides rich semantics for the core parameters: kind (with enum-like values and account numbers), contraAccount (P&L account type), amountMinor (integer Rappen), and periodEnd (balance-sheet date). However, other parameters like costCenterId, sourceRef, and idempotencyKey are not explained, though they are likely self-explanatory or standard.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to describe an Abgrenzung (Swiss accrual) and save it as a draft. It uses a specific verb ('describe' and 'save') and resource ('Abgrenzung'), and the final sentence 'Nothing is posted' distinguishes it from posting-related siblings like accrual_post.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context by explicitly stating this creates a draft and that nothing is posted, implying it is the safe precursor to posting. However, it does not explicitly name alternative tools like accrual_post or accrual_reverse, leaving some inference to the agent. The sibling list includes these related tools, providing additional context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
accrual_discardA
Retire a drafted Abgrenzung before it posts. The row stays on record as discarded with its reason; a posted accrual cannot be discarded (already_posted), reverse it instead.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| accrualId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses that the row is not deleted but remains on record with status `discarded` and a reason, and it names a specific error condition (`already_posted`). It does not mention idempotency effects or reversibility, but the core behavioral profile is clearly communicated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two well-structured sentences with no filler. The primary action and lifecycle constraint are front-loaded, and the error/alternative guidance follows naturally. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the central behavior, state transition, and error case, but it omits important invocation details such as the meaning of `idempotencyKey` and what the response looks like. Since there is no output schema and no annotations, these omissions leave the description slightly incomplete for an agent making a fully informed call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for undocumented parameters. It only adds meaning for `reason` by noting that the retained row keeps the reason. Required parameters like `workspaceId`, `accrualId`, and especially `idempotencyKey` are left completely unexplained, which is a meaningful gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Retire a drafted Abgrenzung before it posts') and clearly identifies the resource. It also differentiates the tool from the reverse operation by explaining that posted accruals cannot be discarded. This makes it immediately distinguishable from siblings like accrual_post and accrual_reverse.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use the tool ('before it posts') and when not to use it ('a posted accrual cannot be discarded'), and names the alternative behavior ('reverse it instead'). This is direct, actionable guidance with no reliance on inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
accrual_getARead-only
Read one Abgrenzung: the row, its lines and reversal lines (from the same function the post uses), and the journal entries it has produced (the accrual, its reversal, the Storno pair).
| Name | Required | Description | Default |
|---|---|---|---|
| accrualId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint=true annotation already declares the read-only safety profile, and the description aligns with it by opening with 'Read'. The description adds meaningful context beyond the annotation by disclosing what is actually returned (row, lines, reversal lines, journal entries, Storno pair) and notes the provenance consistency ('from the same function the post uses'). This is useful behavioral detail for an agent deciding what it can rely on.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one front-loaded sentence with no filler; the verb 'Read' appears immediately and each subsequent clause enumerates a distinct return component. The parentheticals ('from the same function the post uses', 'the accrual, its reversal, the Storno pair') are dense but each earns its place by clarifying provenance or terminology.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read tool backed by readOnlyHint=true, the return contents are well described and the safety profile is covered by annotations. The notable gaps are that both parameters are under-documented across schema and description, and the German term 'Abgrenzung' is used without being fully defined, so the agent needs context from the sibling accusatory set to fully understand the concept being read.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden of explaining the two parameters, but it never mentions workspaceId or accrualId or clarifies their meaning, format, or relationship. The only oblique connection is that 'Read one Abgrenzung' implies the accrualId selects the record; the workspace dimension is entirely unaddressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource — 'Read one Abgrenzung' — and then enumerates the exact return scope: the row, its lines, reversal lines, and produced journal entries (accrual, reversal, Storno pair). This scope is specific enough to clearly distinguish it from the mutating accrual siblings (accrual_create, accrual_post, accrual_reverse, accrual_discard) and from the collection-focused accrual_list. It does not name a sibling, but the read-target and depth are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrasing 'Read one Abgrenzung' implies the intended use — fetch the full detail of a single accrual record — and the listed return components imply the data depth an agent should expect. However, there is no explicit guidance about when to prefer this tool over accrual_list or the read-style siblings (e.g., provision_get), and no exclusions or prerequisites are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
accrual_listARead-only
List Abgrenzungen, newest period first, filtered by periodEnd, status (draft | posted | reversed | discarded) or kind. totalMinor sums the drafts and the posted ones. savedViewId applies a saved view (G00).
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| status | No | ||
| periodEnd | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses meaningful behaviors: results are ordered newest period first, totalMinor aggregates drafts and posted accruals, and savedViewId applies a saved view. This provides useful behavioral context without contradicting the read-only annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences deliver the core action, sorting, filters, status enum, a computed-field behavior, and saved-view semantics. Every phrase earns its place with no redundancy or padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential call dimensions: filtering, ordering, and the special totalMinor/savedView behaviors. It does not document output structure or pagination, but given no output schema exists, but the core information needed to invoke the tool correctly is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by explaining the filter semantics for periodEnd, status (with explicit allowed values), kind, and savedViewId. It does not define periodEnd format or kind value options, but it adds substantial meaning beyond the bare parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), the resource ('Abgrenzungen'/accruals), and concrete dimensions of the operation: newest period first, with filtering by periodEnd, status, and kind. This clearly distinguishes it from singular operations like accrual_get and from related but different list tools such as provision_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context by specifying sort order and available filters, so an agent knows when to call this listing tool rather than a get/create/post operation. It does not explicitly name alternatives or exclusions, but the list-vs-get distinction is strongly implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
accrual_postA
Post a drafted Abgrenzung as an atomic PAIR: the accrual entry dated periodEnd (source accrual) and its automatic Rückbuchung dated the first day after (a real reversal, linked). A locked reversal date rolls the whole pair back, so the books never carry an accrual without the reversal that backs it out. Idempotent per accrual: the same key replays, a different key on a posted accrual is already_posted, a discarded draft is draft_discarded. CONSEQUENCE: Posts the accrual and its automatic next-period reversal as one pair; the only correction afterwards is a reversing pair.
| Name | Required | Description | Default |
|---|---|---|---|
| accrualId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description fully carries the behavioral burden. It discloses the atomic pair, the linked automatic reversal, rollback behavior on locked reversal dates, idempotency semantics, and the specific error outcomes `already_posted` and `draft_discarded`.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and well-structured, with the core purpose front-loaded and consequences summarized. It is somewhat long, but every sentence contributes relevant operational detail for the agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description covers the key behavioral expectations, idempotency cases, and post-condition consequences. It does not specify the success response shape, but for this tool the outcome behavior is well enough described for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage, but the description adds meaningful semantics for `idempotencyKey` (same key replays, different key gives error) and clarifies the operation is on a drafted accrual. It does not explicitly elaborate on `workspaceId` and `accrualId`, but their roles are inferable from the context and naming.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Post') and resource ('drafted Abgrenzung'), and immediately defines the operation as an atomic pair of entries. It distinguishes itself from siblings like accrual_reverse and accrual_discard by emphasizing the posting of a draft and linking to its reversal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It conveys when to use the tool: when a drafted accrual exists and needs to be posted. It also hints at the alternative for correction ('the only correction afterwards is a reversing pair'), though it does not explicitly name a sibling tool like accrual_reverse for that follow-up.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
accrual_reverseA
Revert a posted Abgrenzung without editing history: posts the MIRROR pair (the Storno dated periodEnd, its own reversal dated the day after), so all four entries net to zero on every account in both periods. Refuses already_reversed a second time and not_posted on a draft. CONSEQUENCE: Posts the mirror pair against a posted accrual so every account nets to zero, without rewriting history.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| accrualId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses key behavioral traits: it does not edit history, it posts a mirror pair (Storno dated periodEnd and its own reversal dated the day after), and it refuses certain states. It also warns of the consequence: posting the mirror pair so every account nets to zero. This goes beyond a simple 'reverses an accrual' statement and explains the accounting mechanics. With no annotations provided, the description carries the full burden, and it does so well, though it could mention idempotency behavior or side effects on related records.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-rich, with the core action and consequence front-loaded. The first sentence explains the mechanism and the second adds constraints and a warning. It is slightly redundant: 'posts the MIRROR pair' and 'CONSEQUENCE: Posts the mirror pair' repeat the same action. Trimming the second sentence's repetition would make it tighter, but overall it is concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (accounting reversal with mirror entries, state refusals, no output schema), the description covers the essential behavioral contract: what it does, when it refuses, and the net effect. It does not explain the return value or how to interpret success/failure beyond the refusal states, and it does not detail the `reason` parameter's role. However, for an agent deciding whether to call this tool, the description is largely complete. The lack of output schema and 0% parameter coverage raise the bar, and the description meets it for the core decision but leaves some gaps around parameter specifics and response handling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the four parameters. The description mentions `periodEnd` and `already_reversed`/`not_posted` states, which relate to the accrual being reversed, but it does not explicitly explain the meaning of `workspaceId`, `accrualId`, `idempotencyKey`, or `reason`. The `accrualId` is inferable from the tool's purpose, and `idempotencyKey` is a common pattern, but the description does not add explicit parameter-level semantics. It provides some context (the reversal is dated periodEnd), but not enough to fully compensate for 0% schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: reverting a posted Abgrenzung (accrual) by posting a mirror pair of entries, with specific dates, so all four entries net to zero. It distinguishes itself from related tools like accrual_discard (which likely handles drafts) and provision_reverse by specifying it operates on a posted accrual and refuses drafts. The verb 'Revert' plus the resource 'posted Abgrenzung' is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when the tool can be used: on a posted accrual, and when it cannot: it refuses `already_reversed` a second time and `not_posted` on a draft. This gives clear preconditions and failure modes. However, it does not explicitly name alternative tools for reversing drafts or other reversal scenarios, though the sibling list includes accrual_discard and provision_reverse. The context is clear enough for an agent to know when to invoke this tool, but it lacks explicit 'use X instead' guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
advance_move_stepA
Persist the move-checklist resume point, absolutely (the advance_onboarding_step shape): direction alone starts or re-asserts the move; direction plus step (1-5, done defaulting true) checks a step off or, with done:false, un-checks it (the trial-balance-mismatch recovery); abandon:true deletes the pointer (Umzug abbrechen). Never a gate: the real actions of the journey (create_backup, archive_workspace, the manual legs on the other machine) keep their own validation, and the step timestamps record a claim, not a proof. A different direction than the stored one resets the checklist. completedAt is derived: stamped when all five steps are done, cleared when one is un-done.
| Name | Required | Description | Default |
|---|---|---|---|
| done | No | ||
| step | No | ||
| abandon | No | ||
| direction | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so thoroughly. It reveals that it is not a gate, that timestamps are claims not proof, that completedAt is derived, and that direction change resets the checklist. It also explains the abandon behavior. This is comprehensive for a write operation with multiple modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is detailed but every sentence adds value. It is structured logically: main purpose first, then mode explanations, then clarifications (non-gate, reset, derived timestamp). It is not overly verbose for the complexity of the tool, and the key points are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (multiple modes, interactions with a checklist, derived state) and the absence of an output schema, the description covers everything an agent needs: what it does, how to use each mode, the semantics of step/done/abandon, and the derived completedAt. Nothing critical is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains each parameter's role: direction starts/re-asserts, step (1-5) with done (default true) checks/un-checks, abandon deletes the pointer. It also explains the derived completedAt. This fully covers the parameter semantics beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Persist the move-checklist resume point'. It specifies the resource (move-checklist) and the action (persist/resume point). It also distinguishes itself from the sibling advance_onboarding_step by referencing its shape, making the distinction clear without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong contextual guidance on when to use this tool: it's for the move-checklist, not for validation (explicitly 'Never a gate'). It also describes the various modes (start, check, uncheck, abandon) and how the checklist resets on direction change. It does not explicitly name an alternative tool for onboarding, but the 'advance_onboarding_step' reference implies the separation. Slight gap in naming explicit alternatives, but context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
advance_onboarding_stepA
Persist the first-run wizard resume point: path (fresh|import|demo) and the current step, absolutely, so a closed tab reopens where it left off. Never a business-logic gate: every underlying setup verb keeps its own validation and idempotency, and an agent session skips this bookkeeping entirely (spec US-G03.5). completed:true stamps the wizard finished, once; a later call never un-completes it.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| step | Yes | ||
| completed | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it states persistence is absolute, side effects are limited to bookkeeping, and completed:true is monotonic ('never un-completes it'). This gives an agent a clear mental model of side effects and irreversibility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose, and each clause adds a distinct behavior (persistence, non-gating, completion semantics). The spec reference is compact but provides provenance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 4-parameter bookkeeping tool with no output schema, the description covers the critical invocation details: path enum, step, completed semantics, and side-effect scope. It doesn't specify the allowed format of step or workspaceId, but these are low-risk for a persist operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain the parameters. It enumerates path values (fresh|import|demo), defines step as the current wizard step, and explains the completed flag's irreversible behavior. workspaceId is not explicitly described, though its meaning is conventional.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Persist') and resource ('first-run wizard resume point') and clarifies the exact payload (path, step, completed). It also distinguishes itself from real setup operations by stating 'Never a business-logic gate,' which differentiates it from sibling migration/setup tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit when-not-to-use guidance: 'an agent session skips this bookkeeping entirely' and clarifies it is not a validation or gate mechanism. It doesn't name alternative tools (e.g., get_onboarding_progress), but the exclusions are clear enough to prevent misuse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_askB
Ask the books a question in prose (the Studio composer, D90 D-1). Needs a registered E05 local runtime (needs_local_runtime otherwise); the runtime only classifies the question onto one of the three A26 read models and the executed verb is ALWAYS a read, so no write is reachable from prose. Persists the sentence as a turn (D-5) and records the answering call in the trace.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| periodEnd | No | ||
| periodStart | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden and does more than most: it discloses the runtime requirement, guarantees no write is reachable, and reveals side effects (persisting the turn and recording the call in the trace). It is missing caveats like authorization, failures, or rate limits, but the core behavioral profile is unusually explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-rich, with the primary action front-loaded and no filler sentences. The internal code names and parentheticals reduce clarity, but every clause carries either a constraint, a side effect, or a purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with no annotations, no output schema, and no parameter descriptions, the definition is incomplete: it does not state what the answer looks like, how period filters interact with the question, or how idempotency applies. It provides strong behavioral context but leaves invocation semantics underspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only obliquely covers 'text' as the prose question. It never explains workspaceId, periodStart, periodEnd, or idempotencyKey, which an agent needs to invoke the tool correctly. The description does not compensate for the schema's lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a verb and resource: 'Ask the books a question in prose,' and adds that the verb is always a read, distinguishing it from write-oriented sibling tools. However, the meaning of internal codes like 'D90 D-1' and 'A26 read models' is opaque, and the description never explicitly says what the tool returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a clear prerequisite ('Needs a registered E05 local runtime') and implies the intended use is natural-language questions over the books. But it gives no explicit guidance about when not to use this tool or which sibling to prefer, e.g., ledger_qa or detect_anomalies, so routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_prose_deleteB
Delete one agent session’s stored prose (D90 D-5): every turn text of the session is cleared, the call trace stays. The words are the user’s to remove; what actually ran is the trust view’s substrate and remains. Owner-gated (manage_agent_dial); idempotent per key.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of disclosing behavioral traits. It reveals important details: clears all turn text, preserves call trace, requires owner-gate (manage_agent_dial), and is idempotent. However, it doesn't state what happens to the session itself or whether deletion is reversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is concise, with key information front-loaded: the action, scope, and implications are stated in the first sentence. Later sentences add ownership and idempotency details without redundancy. Could be slightly more structured but is well within appropriate length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a simple deletion tool but lacks essential parameter semantics and does not clarify the return value or error states. Given the tool's complexity (3 params, no output schema), it could provide more detail on what the agent should expect after deletion.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description provides no explanation of the parameters (workspaceId, sessionId, idempotencyKey). It only mentions the idempotency key indirectly by saying 'idempotent per key,' but gives no format or usage details for any parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool deletes a specific agent session's stored prose and specifies the scope (every turn text cleared, call trace remains). It distinguishes itself from sibling tools like agent_trust_summary and get_agent_session by focusing on deletion, though it doesn't explicitly name a sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides context on when to use (owner-gated, idempotent per key) and implies it's for removing user-authored text while preserving the audit trail. However, it doesn't explicitly state when not to use it or mention alternative tools for related operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_trust_summaryARead-only
The trust evidence per dial capability (Vertrauen): stored and effective level, who set it and when, and the proposed/approved/rejected/auto-executed counts over a stated window (default the last 90 days). Derived from the trace at read time; never stored, never a score.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description goes further by stating the data is 'derived from the trace at read time; never stored, never a score.' This adds meaningful behavioral context beyond the annotation, preventing misconceptions about persistence and semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two focused sentences convey the resource, the exact contents, the time-window behavior, and the key semantic caveat ('never a score'). No filler or repetition; the most important qualifiers are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description enumerates the main return contents and clearly states the computation and persistence model. It lacks exact response structure and date format details, but for a simple read-only evidence tool, it is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It adds the concept of a 'stated window (default the last 90 days)', giving some meaning to the from/to parameters, but it does not name them, specify date formats, or clarify the required workspaceId. The compensation is partial but not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific deliverable ('trust evidence per dial capability'), enumerates the data it returns (stored/effective levels, who set it, when, counts over window), and explicitly distinguishes it from a score. This differentiates it from sibling tools like get_agent_dial and set_agent_dial.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: this is a read-only, trace-derived evidence view over a default 90-day window. It implies when to use it (when historical trust evidence is needed) without explicitly naming alternatives, but the context is sufficient for an agent to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
aging_reportARead-only
The receivables summary behind the aging bar: the same open items list_open_items itemises, totalled per bucket and per customer, largest debtor first. byBucket is the invoice-currency sum and baseByBucket the base-currency one, which is the figure that ties to baseTotalOpenMinor in a workspace holding more than one currency. Use it to answer "how much is over 90 days overdue and who owes it" in one call. Carries the same reconciliation flag against 1100 Debitoren. Reads only.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations provide readOnlyHint=true, and the description reinforces this with 'Reads only.' It adds useful behavioral context beyond that: the relationship to baseTotalOpenMinor in multi-currency workspaces conveniently ties to reconciliation flags against account 1100 Debitoren. No contradiction exists between description and annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient: each sentence contributes meaningful information (report nature, aggregation method, currency handling, use case, reconciliation flag, read-only status). It is front-loaded with the core summary, though it could be slightly trimmed without losing value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description does a decent job of explaining the report's output shape: totals per bucket and customer, and the difference between byBucket and baseByBucket. However, it does not mention the asOf parameter or any pagination/limits for large result sets, leaving some operational context missing. For a read-only report, this is a moderate gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, providing only parameter names (workspaceId and asOf) without any meaning. The description does not explain these parameters at all; it only mentions output-related fields like byBucket and baseByBucket. The agent must guess what asOf does (likely a date for the aging snapshot), and workspaceId is presumed from the required flag. This is a clear gap in compensating for the schema's lack of detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states what the tool does: it provides a receivables summary totalled per aging bucket and per customer, with the largest debtor first. It explicitly differentiates itself from list_open_items by noting it is the same open items 'totalled' rather than itemised, and gives a concrete use case ('how much is over 90 days overdue and who owes it').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear usage context: 'Use it to answer ... in one call.' It also references list_open_items as the underlying itemised view, implying the alternative for detailed line items. However, it does not explicitly state when not to use this tool or mention other alternatives, so it lacks exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
allocate_paymentA
Allocate a parked Guthaben to open items. The money was already booked when the payment arrived, so this posts no entry: it records which open items that money settles. Requires intent='allocate_payment'. A wrong allocation is corrected by reversing the whole payment and recording it again. CONSEQUENCE: Applies a recorded payment to open items; the settlement posts to the ledger.
| Name | Required | Description | Default |
|---|---|---|---|
| intent | Yes | ||
| paymentId | Yes | ||
| allocations | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full burden. It discloses the critical side effects: no posting at allocation time, the record of which open items are settled, and that the settlement posts to the ledger. It also explains how a wrong allocation is corrected, which is important for a consequential mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and kept to a few sentences; the correction and consequence details earn their place. The CONSEQUENCE line somewhat restates the opening sentence, adding minor redundancy but no significant bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the prerequisite (payment already booked), required intent, ledger side effect, and correction method, which are the hard-to-infer parts. It does not describe the allocation object fields or the return/idempotency behavior, and there is no output schema, but it is reasonably complete for a well-scoped action.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Adds meaning to intent (exact required value), paymentId (the already-recorded payment), and allocations (they target open items). However, with 0% schema description coverage, it leaves the allocation subfields (skontoMinor, writeOffMinor, targetKind, etc.) and idempotencyKey semantics largely unexplained, so compensation is only partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('Allocate') on a concrete resource ('parked Guthaben'/'recorded payment') to open items, and clarifies that it posts no new payment entry because the money was already booked. This semantically distinguishes it from sibling payment-recording tools despite no sibling being named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: this tool is for payments already booked when they arrived, requires intent='allocate_payment', and gives the correction path (reverse the whole payment and record it again). It does not explicitly name alternatives or state when not to use it, so it stops short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
apply_qr_matchA
Settle an invoice from a queued credit, delegating the posting to A14 record_payment (source 'qr'): debit Bank, credit Debitoren, document status and Ist-VAT stamp included, A21 posts nothing itself. Every figure is NET of linked credit notes: a payer holding a Gutschrift owes the principal, and paying it exactly settles in full. mode defaults to 'partial' (allocates up to the open amount, parks any surplus as Guthaben, never writes off); mode 'full' additionally writes off a SHORT credit's residual, and only within the A14 one-click write-off threshold (default CHF 1.00): a larger residual refuses with write_off_above_threshold naming the amount (type a deliberate larger Ausbuchung through record_payment). A foreign-currency credit refuses with currency_mismatch. P8: pass confirmed=true, or the auto-apply dial must be ON with a LIVE 'high' score for exactly this invoice. NOT automatable (D77): a stored rule may never make this judgment. Idempotent per credit AND per decision: the same invoice+mode replays, a different mode refuses honestly as already_applied. CONSEQUENCE: Settles the invoice from the queued credit and posts the payment; the settlement can only be undone by a reversal.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | ||
| creditId | Yes | ||
| confirmed | No | ||
| invoiceId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so thoroughly. It discloses the posting delegation, net-of-credit-note behavior, mode defaults, write-off threshold, exact error outcomes (write_off_above_threshold, currency_mismatch, already_applied), idempotency guarantees, and the P8 confirmation requirement. It even states the settlement is only undoable by reversal and that a stored rule may never make this judgment, which is far beyond the minimum.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long and dense, but every paragraph addresses a distinct aspect: mechanism, netting, mode semantics, safeguards, idempotency, and consequence. It is front-loaded with the action verb and structured with labels like 'CONSEQUENCE'. Minor redundancy exists between the opening action and the consequence statement, but overall the length is justified by the high-risk domain.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and lack of an output schema, the description covers the main outcomes, failure modes, and automation constraints well. It names specific error identifiers and explains the operational consequence. It does not describe the expected return payload or required permissions, but for selection and invocation correctness this is largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage, but the description compensates well for the most ambiguous parameters: mode is fully explained (partial vs full, write-off threshold), confirmed is connected to the P8 rule, and idempotencyKey semantics are implied by the idempotency statement. It does not explicitly define creditId, invoiceId, or workspaceId, though these are inferable from context, so the compensation is strong but not exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Settle an invoice from a queued credit', and immediately distinguishes the tool from record_payment by stating it delegates posting to A14 record_payment (source 'qr') and that A21 posts nothing itself. It is clear which object the tool acts on and how it relates to the payment pipeline, setting it apart from nearby siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit conditions for mode usage, states the alternative path for larger write-offs ('type a deliberate larger Ausbuchung through record_payment'), and includes a firm automation restriction ('NOT automatable (D77)'). However, it does not directly contrast with closely related siblings like match_qr_payment or override_qr_match, so an agent is not fully told which tool to choose among those.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
approve_drafted_actionA
Approve a drafted agent action from the inbox: replay its verb through the shared dispatch as the approver (RBAC re-checked, idempotent on the stored key, so approve-twice never double-posts) and mark it executed. allowFuture additionally records the standing per-capability grant to auto (D103): the same attributed, revocable dial write set_agent_dial performs. The agent can never self-approve; a second, independent actor must review. Owner-gated (manage_agent_dial).
| Name | Required | Description | Default |
|---|---|---|---|
| actionId | Yes | ||
| allowFuture | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries the burden of behavioral disclosure. It explicitly mentions idempotency ('approve-twice never double-posts'), RBAC re-checking, the side effect of allowFuture (recording a standing grant via set_agent_dial), the self-approval prohibition, and the owner-gating requirement. These are significant behavioral traits beyond what the schema reveals.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph but avoids fluff. It front-loads the primary purpose, then adds crucial behavioral details. Every clause contributes information about idempotency, permissions, side effects, or constraints. It is appropriately sized for a tool with three parameters and no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (3 params, no output schema, no annotations), the description covers all essential aspects: what it does, how it does it, idempotency, permission requirements, side effects of allowFuture, self-approval restriction, and owner gating. An agent has enough information to decide when and how to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must clarify parameters. It explicitly explains allowFuture: 'records the standing per-capability grant to auto (D103): the same attributed, revocable dial write set_agent_dial performs.' workspaceId and actionId are inferable from the action context ('from the inbox', 'drafted agent action'), though not explicitly defined. This adds meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the precise action: 'Approve a drafted agent action from the inbox' with a clear verb and resource. It further explains the mechanism ('replay its verb through the shared dispatch as the approver') and distinguishes it from sibling actions like reject_drafted_action and list_drafted_actions by focusing on approval semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides usage context: 'The agent can never self-approve; a second, independent actor must review' and 'Owner-gated (manage_agent_dial)' indicate when and by whom the tool should be used. It does not explicitly contrast with alternatives, but the purpose is clear enough that an agent can infer when to use this tool versus rejection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
approve_entryA
Sign off one checked entry (freigeben): sets its review status to approved, as metadata, neither locking nor altering the entry. Approving an already-reversed entry is allowed and reported with alreadyReversed. A sign-off is a human act: requires the review capability (Treuhänder/owner) and no automation rule may fire it.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| entryId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It clearly states the operation is metadata-only, does not lock or alter the entry, handles already-reversed entries with an alreadyReversed report, and imposes a human-approval constraint. It lacks response-format details but covers the most important behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler: the core effect is front-loaded, the edge case follows, and the permission constraint closes the description. Each sentence contributes meaningful decision-relevant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is behaviorally rich but incomplete for invocation: the required idempotencyKey is unexplained, the note parameter is ignored, and there is no output schema to clarify the response or alreadyReversed return shape. Given four parameters and zero schema descriptions, an agent would still need external knowledge to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not compensate. It only refers to 'one checked entry' but never explains the workspaceId, entryId, idempotencyKey, or note parameters. An agent cannot infer idempotency semantics or which parameter identifies the entry from either schema or description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Sign off one checked entry') and its exact effect ('sets its review status to approved'), so an agent knows precisely what the tool does. It also distinguishes this from sibling approval tools like approve_drafted_action and expense_claim_approve by clarifying it targets checked entries and is metadata-only.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: it is a human sign-off, requires the review capability (Treuhänder/owner), and must not be fired by automation rules. It does not explicitly name alternatives or state when not to use this tool, but the permission and human-act constraints effectively guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
archive_accountB
Archive an account (soft; the account keeps its postings).
| Name | Required | Description | Default |
|---|---|---|---|
| accountId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It discloses that the operation is soft and preserves postings, which is a meaningful side effect. However, it does not mention reversibility, what happens to future postings, or visibility effects, leaving gaps in the behavioral picture.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler. The core action is front-loaded and the critical qualifier '(soft; the account keeps its postings)' is placed immediately, giving the agent the most important behavioral fact without extra words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool, the description covers the essential purpose and the most important side effect. It is minimally viable, but it omits contextual hints such as reversibility via unarchive_account or implications for postings/reporting, which would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate for the undocumented parameters. It adds no information about workspaceId or accountId beyond their obvious names. Since the parameter names are self-explanatory, this is not a complete failure, but the description provides zero semantic addition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Archive') and resource ('account'), and the parenthetical '(soft; the account keeps its postings)' distinguishes it from a hard delete. It is clear enough to separate from delete_account, though it does not explicitly contrast with unarchive_account.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives like unarchive_account or delete_account. The 'soft' hint implies a non-destructive use case, but there is no explicit when-to-use or when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
archive_automation_ruleA
Retire a rule: it stops matching, disappears from the default list, and stops being enable-able. Never a delete, and its run history is untouched, because a firing that happened stays true after the rule that caused it is gone.
| Name | Required | Description | Default |
|---|---|---|---|
| ruleId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it thoroughly discloses the operation's safety profile: 'Never a delete' and 'run history is untouched' clarify both irreversibility and data preservation. It also explains the behavioral side effects beyond just saying 'archive,' going well beyond what the schema could convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action and its main effects. The prose is efficient and meaningful; the 'firing that happened stays true' phrase adds color but reinforces an important guarantee, so the extra length still earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, the description covers what changes, what does not change, and the non-delete nature, which is enough to invoke it correctly in most cases. It omits explicit return or error behavior and parameter details, but the tool's simplicity makes those omissions less critical.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Both parameters are bare strings with zero schema descriptions (0% coverage), and the description adds no guidance on their semantics or format. The names 'workspaceId' and 'ruleId' are self-explanatory, but the description does not compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with 'Retire a rule,' naming a clear verb and resource, then spells out three concrete effects: stops matching, disappears from the default list, and stops being enable-able. This clearly distinguishes it from enable, disable, and delete semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides useful context by stating the permanent consequence ('stops being enable-able') and explicitly ruling out delete, which helps an agent infer this is the terminal option. However, it never names sibling tools such as disable_automation_rule or states when to choose that alternative instead, so the selection guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
archive_bank_accountA
Archivieren: hide a closed Bankkonto from pickers while keeping it referenceable by historical statements and entries. Never deletes the account.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| bankAccountId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavior disclosure. It reveals the key behavioral traits: hides from pickers, retains historical references, and never deletes. This goes well beyond the schema but doesn't mention idempotency, reversibility, or effects on existing picker selections.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the verb and core purpose, then adds the critical non-deletion guarantee. No waste; every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential behavior but omits prerequisites (e.g., account must be closed), what the response contains, and the connection to unarchive_bank_account. With no annotations and no output schema, more operational detail would be expected for full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain workspaceId, bankAccountId, or idempotencyKey. While the names suggest their roles, the description adds no parameter-level guidance, so it fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (hide a closed bank account from pickers) and resource (Bankkonto), and clarifies the key distinction from deletion: 'Never deletes the account.' This clearly separates it from delete_account and other destructive archive tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear context: use for closed bank accounts that should be hidden from pickers but remain referenceable. It does not explicitly name alternatives like unarchive_bank_account or state when not to use it, which prevents a 5, but the context is specific enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
archive_contactC
Archive a contact (soft).
| Name | Required | Description | Default |
|---|---|---|---|
| contactId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. The parenthetical 'soft' is genuine behavioral information: it indicates the contact is not permanently destroyed and can likely be restored. However, it does not disclose side effects, permissions, idempotency, or what happens to associated data, so transparency is only partially addressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with no filler words, and the core action is front-loaded. It is appropriately short for a simple two-parameter tool. However, the brevity comes at the cost of missing usage and behavioral context, so it stops just short of an exemplary score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description is too sparse to be considered complete. It says nothing about the result of archiving, whether the contact disappears from lists, or how the operation can be reversed. The agent can guess the basics from the name and the word 'soft,' but meaningful context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the lack of parameter documentation, but it does not mention either parameter. The names workspaceId and contactId are somewhat self-explanatory, yet the description adds no meaning beyond the raw schema. This is a clear gap given the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Archive a contact.' The parenthetical 'soft' adds important meaning that this is not a hard delete, which helps distinguish it from deletion-style tools. It does not explicitly name sibling tools like unarchive_contact, but the core purpose is still clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as unarchive_contact, update_contact, or other archive operations. The word 'soft' implies a reversible action, but the description does not state when archiving is appropriate or how to reverse it. This leaves usage decisions entirely to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
archive_cost_centerC
Archive a cost centre (soft).
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| costCenterId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Archive a cost centre (soft)' implies a mutation but does not say whether it is reversible, what the soft-archive state means, whether it prevents editing or deleting, who can perform it, or what happens to associated data. The word 'soft' hints at reversal but is not elaborated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short sentence that is easy to parse and front-loads the verb and object. The parenthetical '(soft)' is a useful qualifier, though it could be more explicit. No wasted words, but the conciseness comes at the cost of missing necessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations, no output schema, and 0% parameter documentation, this description is incomplete. It lacks prerequisites (e.g., existence of the cost center), what happens after archiving, whether the operation is reversible, and parameter meaning. The sibling set includes lifecycle tools (create_cost_center, unarchive_cost_center, delete_cost_center, list_cost_centers), but the description does not situate this tool within that lifecycle.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the two parameters (workspaceId, costCenterId) are entirely undocumented. The description names neither parameter and provides no hint about formats, scoping, or dependencies. An agent must guess which workspace and cost center identifiers are expected.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Archive a cost centre (soft).' clearly states the action (archive) and the resource (cost centre), and the parenthetical '(soft)' usefully indicates this is a soft archive (reversible, not a hard delete). It is distinguishable from the sibling tools 'delete_cost_center' and 'unarchive_cost_center'. However, it doesn't explicitly state that this is a status change or describe the state transition, which would add clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use this tool versus alternatives like 'delete_cost_center' or 'unarchive_cost_center'. The '(soft)' hint implies a contrast with hard delete, but it doesn't state the relationship or offer decision criteria. An agent would have to infer usage from the name and sibling proximity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
archive_document_templateA
Soft-archive a template (and clear it as default). NOT destructive: a template frozen onto an already-issued document keeps rendering that document forever; archiving only removes it from the active list and from new defaults.
| Name | Required | Description | Default |
|---|---|---|---|
| templateId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden, and it does so excellently. It discloses that the operation is non-destructive, that templates frozen onto issued documents continue rendering forever, and that archiving only affects active availability and default selection. This is precisely the kind of behavioral detail an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight sentences with no filler. The core action is front-loaded, and the clarifying non-destructive detail immediately follows in a parenthetical that earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two required simple parameters and no output schema, the description covers the action, the side effect, the non-destructive guarantee, and the ongoing impact on issued documents. The only minor gap is that it does not mention the purpose of idempotencyKey, but that is a standard parameter and not critical for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate for the parameters' meaning, but it does not explain workspaceId, templateId, or idempotencyKey beyond referencing 'a template' in general terms. The parameter names are somewhat self-explanatory, but the description adds no semantic value to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('soft-archive'), the resource ('a template'), and the side effect ('clear it as default'), making the tool's purpose immediately clear. It also distinguishes itself from destructive operations by explicitly saying 'NOT destructive', which is meaningful given the sibling set contains many archive/delete tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the practical consequence—archiving removes the template from the active list and from new defaults while preserving rendering on already-issued documents—so an agent can judge when this is appropriate. It does not explicitly contrast with a delete/unarchive alternative, but there is no obvious sibling that would compete for the same call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
archive_fieldA
Archive a custom field so it stops being offered as an input. Never a delete: every value already stored survives and stays readable through list_field_values. Returns how many saved views still name the field, so the caller can warn rather than be surprised later.
| Name | Required | Description | Default |
|---|---|---|---|
| fieldDefId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it reveals that archiving is non-destructive, that values remain readable, and that the response includes a count of saved views still referencing the field. It does not mention reversibility or permissions, but the most decision-relevant behaviors are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: the main purpose comes first, followed by the critical non-destructive guarantee, then the useful return-value behavior. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description covers the essential call context: what it does, what it does not do, where data remains readable, and what the caller should do with the response. The only notable gap is a lack of parameter-level detail, but the required IDs are reasonably inferable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it does not explain workspaceId, fieldDefId, or idempotencyKey. It only implies fieldDefId identifies the custom field. The parameter names are self-explanatory to a degree, but the description adds little semantic value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Archive'), a specific resource ('a custom field'), and the concrete effect ('stops being offered as an input'). It also draws a clear line against deletion, which distinguishes this from destructive sibling operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear contextual guidance: use this when you want to stop offering a field while preserving stored values. The explicit 'Never a delete' warning and the pointer to list_field_values provide useful boundaries, though it does not name a specific non-archive alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
archive_itemC
Archive an item (soft).
| Name | Required | Description | Default |
|---|---|---|---|
| itemId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of disclosing behavior. The word 'soft' hints that the operation is non-destructive, but it does not explain side effects (e.g., whether the item disappears from listing, whether it can be restored, what happens to related entities). For a mutation tool, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, short sentence with no redundant words, which is concise. However, its brevity results in under-specification; it is not structured to convey key information. It is minimal but incomplete, so it earns a middle score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two required parametersrande no annotations or output schema, the description is far too sparse. It omits the action semantics (what 'soft' means), the effect on item availability, any permission requirements, and the expected response. An agent cannot reliably invoke this tool correctly based solely on the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the lack of parameter documentation. It does not mention workspaceId or itemId at all. While the parameter names are somewhat self-explanatory, the description adds no context, semantics, or format expectations, leaving the agent to infer complex requirements like workspace context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Archive') and resource ('item'), and the parenthetical '(soft)' signals that this is a non-destructive operation, distinguishing it from a hard delete like delete_item. However, it does not explicitly contrast with unarchive_item or define what 'soft' means in this domain, so the clarity is good but not fully comprehensive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives like unarchive_item, delete_item, or archive_account. The description provides no context on prerequisites, reversibility, or selection criteria. An agent is left to infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
archive_roleA
Soft-flag a custom role so it stops being offered. Never a delete, so historical attributions still resolve, and refused while any member still holds it (role_in_use).
| Name | Required | Description | Default |
|---|---|---|---|
| roleId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses that the operation is non-destructive, that it keeps historical attributions resolving, and that it fails with 'role_in_use' when the role is still assigned. This is strong transparency for a soft-delete operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences front-load the core behavior and then add the critical safety distinctions. Every word earns its place, with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a straightforward soft-archive operation, the description covers the essential behavioral semantics, the key failure condition, and the non-delete nature. It does not address idempotency behavior or reversal, but these are secondary for selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not compensate by explaining the parameters. 'roleId' and 'workspaceId' are self-explanatory from their names, but 'idempotencyKey' is left completely unexplained, and no parameter-level guidance is added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('Soft-flag'), the resource ('a custom role'), and the observable effect ('stops being offered'). It explicitly differentiates itself from a delete operation, which is essential for tool selection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context: this is not a delete, it preserves historical attributions, and it will be refused while any member holds the role. It stops short of naming an explicit alternative tool, but the conditions and exclusions are well stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
archive_workspaceA
Archivieren: retire a finished mandate (archived:true) or put it back in play (archived:false). A reversible flag, never a delete: the books stay intact and exportable (OR 958f), reads keep answering, and every OTHER write into an archived workspace returns workspace_archived.
| Name | Required | Description | Default |
|---|---|---|---|
| archived | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden and fully meets it: it declares reversibility ('a reversible flag'), data preservation ('the books stay intact and exportable'), continued read access ('reads keep answering'), and the workspace_archived error on other writes. This gives an agent an accurate behavioral model without needing to test the tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the purpose, and covering side effects economically. The only waste is the 'Archivieren:' prefix, which duplicates the tool name, and the cryptic '(OR 958f)' reference that may read as noise for an agent. Overall the density of useful information per sentence is high.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param tool with no output schema and no annotations, the description covers purpose, both flag directions, data preservation, and blocked-write behavior — enough to call correctly. The remaining gaps are the return value (unknown, since no output schema exists) and the semantics of idempotencyKey, which are minor for a simple toggle operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates for the key semantic parameter: it defines both archived:true and archived:false outcomes in depth, which is the parameter carrying real meaning. workspaceId is self-evident from the tool name, but idempotencyKey's purpose (why it is required, what value to supply) is left unexplained — a real but minor gap given its conventional meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with the specific action — 'retire a finished mandate (archived:true) or put it back in play (archived:false)' — giving a clear verb, resource, and both flag states. The workspace scoping and the 'never a delete' clause distinguish it from the many archive_account/archive_contact/archive_item siblings and destructive delete tools. Nothing about the purpose is ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names the condition to use the tool: a mandate is finished (archive) or needs reactivation (unarchive). The 'never a delete' statement is an explicit when-not, telling the agent this is not the tool for permanent data removal. However, it stops short of naming a specific sibling alternative for permanent deletion (e.g., env_delete, delete_account), so guidance is context-rich but not fully explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_acquireA
Record the PRIMARY acquisition of a draft fixed asset (H02): the first financial event, which capitalises the asset. In ONE atomic transaction it writes a sub-ledger asset_transaction (type=acquisition) AND posts the balanced GL journal via A02 (Dr the asset GL account inherited from the category, Cr the chosen creditAccountId, for acquisitionCostRappen), then moves the asset draft -> active and confirms its cost base. creditAccountId must be a real asset/liability/equity account (credit_account_wrong_type otherwise; invalid_credit_account when missing or foreign). An optional residualValueRappen (0..cost) overrides the residual; optional costCenterId stamps both legs; optional source (manual|vendor_bill|project|opening) + sourceDocumentId link the event (source_document_not_found if the id does not resolve). Refused with invalid_cost (cost <= 0), asset_not_acquirable (disposed/archived), already_acquired (a primary acquisition already exists), invalid_residual, or period_locked (date in a hard-locked A03 period). After success the H01 financial-field lock is active. Idempotent on idempotencyKey: a replay returns the original {asset, transaction, journalEntry} and posts no second entry. CONSEQUENCE: Capitalises the asset and posts its acquisition entry; the financial fields lock afterwards.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | ||
| source | No | ||
| assetId | Yes | ||
| description | No | ||
| workspaceId | Yes | ||
| costCenterId | No | ||
| idempotencyKey | Yes | ||
| creditAccountId | Yes | ||
| sourceDocumentId | No | ||
| residualValueRappen | No | ||
| acquisitionCostRappen | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since no annotations are provided, the description carries the full burden. It thoroughly discloses behavior: atomic transaction, GL posting, state transitions, idempotency on idempotencyKey, consequences (financial-field lock), and error conditions. It adds rich context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but information-dense, with each sentence adding value. It front-loads the primary purpose and then lists constraints and consequences. Slightly overlong but justified given the complexity of the operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the high parameter count (11) and zero schema coverage, the description is remarkably complete. It covers all major aspects: atomicity, accounting effect, validation errors, idempotency, and post-conditions. An agent can confidently invoke this tool without further research.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains several parameters: creditAccountId must be a real account, residualValueRappen within 0..cost, source and sourceDocumentId linkage, idempotencyKey for replay. It does not explain all 11 parameters but covers the critical ones, leaving a minor gap for date/workspaceId/etc.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool records the PRIMARY acquisition of a draft fixed asset (H02), the first financial event that capitalises the asset. It specifies the exact verb, resource, and unique purpose, distinguishing it from siblings like asset_add_capitalisation or asset_dispose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly mentions when to use this tool (for primary acquisition, when the asset is in draft state) and lists refusal conditions (e.g., asset_not_acquirable, already_acquired). However, it does not explicitly contrast with alternatives like asset_add_capitalisation, though the context of 'primary acquisition' implies it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_acquisition_summaryARead-only
What was capitalised in a period (H09, US-H09.5): every acquisition and additional-capitalisation event (H02) posted between fromDate and toDate (ISO, inclusive) with its asset number/name, acquisitionDate, costRappen (the capitalised delta), category code/name, the GL asset account (id/number/name), the source (source_document_type or "manual"), sourceDocumentId and journalEntryId. Footer totals: count and sum cost. Optional filter by categoryIds and locationIds. Empty range is an empty success. Pure read, §H-TENANT.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No | ||
| toDate | Yes | ||
| fromDate | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and the description aligns with it by stating 'Pure read, §H-TENANT'. Beyond the annotation it adds valuable behavioral context: ISO dates are inclusive, an empty range is an empty success, footer totals are returned, and the source field covers source_document_type or 'manual'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A dense single sentence enumerating output fields, followed by two short sentences for totals, empty-range behavior, and read-only status. Front-loaded with the purpose; every element earns its place, though the field enumeration makes it long.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Since there is no output schema, the description carries the full burden of describing the return shape, and it does so thoroughly: all fields, footer totals, filters, and the empty-range success behavior. Missing only minor details such as result ordering and pagination limits.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description must compensate. It adds meaning for fromDate/toDate (ISO, inclusive) and filter.categoryIds/locationIds (optional), but the required workspaceId parameter is never mentioned, leaving a gap for a required input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete resource and scope: 'What was capitalised in a period (H09, US-H09.5)' covering acquisition and additional-capitalisation events (H02). It lists output fields explicitly and is clearly distinguishable from siblings like asset_acquire (a mutation), asset_transaction_list (all transactions), and asset_depreciation_* (depreciation).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage context: it is the H09/H09.5 capitalisation report, a pure read, with date-range and optional category/location filters. It does not explicitly name alternatives or state when-not-to-use, but the scope definition is precise enough to route an agent correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_add_capitalisationA
Capitalise ADDITIONAL cost onto an already-acquired asset (H02): an improvement or major overhaul. In ONE atomic transaction it writes a sub-ledger asset_transaction (type=additional_capitalisation) AND posts the same balanced Dr asset / Cr creditAccountId pattern via A02 for amountRappen, then raises the asset acquisition_cost_rappen and net_book_value_rappen by that amount (accumulated depreciation is left unchanged; residual and useful life are untouched). Requires a live asset carrying a primary acquisition (not_acquired otherwise; asset_not_acquirable when disposed/archived). Same credit-account, source-document and period rules as asset_acquire. Idempotent on idempotencyKey. CONSEQUENCE: Posts additional cost onto an acquired asset and raises its book value.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | ||
| source | No | ||
| assetId | Yes | ||
| description | No | ||
| workspaceId | Yes | ||
| amountRappen | Yes | ||
| costCenterId | No | ||
| idempotencyKey | Yes | ||
| creditAccountId | Yes | ||
| sourceDocumentId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description fully carries the behavioral disclosure burden. It spells out the atomic transaction, the exact sub-ledger entry type, the posting pattern, the effect on acquisition_cost_rappen and net_book_value_rappen, what stays unchanged (accumulated depreciation, residual, useful life), and idempotency on idempotencyKey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and well-structured: definition first, then atomic mechanics, prerequisites, shared rules, idempotency, and consequence. The final CONSEQUENCE sentence is mildly redundant with the opening, but the overall density and organization still make this an efficient, valuable description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex multi-step mutation with no annotations and no output schema, the description is unusually thorough: it covers effects, unchanged fields, prerequisites, error conditions, and idempotency. It is not fully complete because parameter-level semantics are only partially described and no return/output information is given, but the critical behavioral contract is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. It does explain amountRappen and creditAccountId via the posting pattern and idempotencyKey via idempotency, and it alludes to shared source-document rules. However, it leaves workspaceId, assetId, date, source, description, costCenterId, and sourceDocumentId without explicit semantics, so not all parameters are fully clarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and object: 'Capitalise ADDITIONAL cost onto an already-acquired asset (H02): an improvement or major overhaul.' This clearly distinguishes the tool from initial acquisition tools like asset_acquire and from disposal/transfer tools, making its purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states prerequisites ('Requires a live asset carrying a primary acquisition'), failure conditions ('not_acquired otherwise; asset_not_acquirable when disposed/archived'), idempotency behavior, and refers to the sibling asset_acquire for shared credit-account, source-document and period rules. This gives an agent concrete when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_archiveA
Soft-archive an asset (status=archived): it leaves the active register but stays readable. Refused with asset_in_use while the asset is active or fully_depreciated (an asset the ledger still depends on cannot be hidden; dispose it via H06 instead). Only a draft or already-terminal asset may be archived.
| Name | Required | Description | Default |
|---|---|---|---|
| assetId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden, and it does substantial work: it discloses soft-archive semantics (leaves register, stays readable), the specific error code (asset_in_use), the rationale ('an asset the ledger still depends on cannot be hidden'), and valid starting states. It does not disclose whether the operation is reversible or what the response contains, leaving minor residual gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed sentences with zero filler; the core action and its most important consequence are front-loaded, and the error condition, alternative, and precondition each earn their place. The H06 reference is the only slightly cryptic element, but it is not redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with no output schema and no annotations, the description covers behavior and failure modes well, but two gaps remain: the alias 'H06' is unexplained jargon an agent may not resolve to a specific sibling, and the required idempotencyKey's role is never addressed. Reversibility (unarchive path) is also unstated, leaving the picture adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the three required string parameters (workspaceId, assetId, idempotencyKey), but it only clarifies the asset's semantic role implicitly via the tool's purpose. It adds nothing about the meaning or constraints of idempotencyKey — a required parameter — or how workspaceId scopes the operation, so it fails to compensate for the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description leads with a precise verb+resource+state-change: 'Soft-archive an asset (status=archived)' and immediately differentiates itself from hard deletion: 'it leaves the active register but stays readable.' This clearly distinguishes it from sibling tools like asset_dispose, archive_item, and asset_update without requiring the agent to open any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit when-to-use guidance by naming the error condition ('Refused with asset_in_use while the asset is active or fully_depreciated'), naming the alternative ('dispose it via H06 instead'), and stating the precondition ('Only a draft or already-terminal asset may be archived'). An agent knows exactly when to call this versus a disposal path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_category_archiveA
Soft-archive a category (active=false): it leaves the default picker but stays resolvable for historical assets. Refused with category_in_use while any asset (even disposed) still references it. Deletion is never offered (a category must stay resolvable for its assets whole depreciable life).
| Name | Required | Description | Default |
|---|---|---|---|
| categoryId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral disclosure. It explains the exact effects (active=false, leaves default picker, stays resolvable), the failure mode (refused with category_in_use while any asset references it), and the rationale (category must stay resolvable for its depreciable life). This gives the agent a complete picture of side effects and constraints without needing annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each carrying essential information: the core action, the error condition, and the rationale for no deletion. It is front-loaded with the action and avoids any fluff or repetition, making it highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description covers all critical aspects: what the tool does, side effects, failure conditions, and why deletion is not offered. It even explains the 'soft-archive' nuance (leaves default picker, stays resolvable). Nothing essential is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage for its three parameters (categoryId, workspaceId, idempotencyKey), and the description does not mention any of them. The parameter names are somewhat self-explanatory (e.g., categoryId identifies the category), but the description adds no additional meaning about their purpose, formats, or why idempotencyKey is required. Given the low schema coverage, the description should compensate but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool soft-archives a category by setting active=false, and contrasts it with deletion. It distinguishes the tool from other category operations like asset_category_update or asset_category_delete (if it existed) by explaining the soft-archive semantics, making it unambiguous what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when this tool is appropriate: when you need to soft-archive a category that is not referenced by any asset. It also explains the error condition (category_in_use) and states that deletion is never offered, which implicitly guides the agent away from a delete operation. However, it does not explicitly name a sibling alternative (e.g., 'use asset_category_delete for hard delete'), so it lacks an explicit when-not-to-use reference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_category_createA
Create a fixed-asset category carrying the defaults every asset created under it inherits (H01): a depreciation method (straight_line | declining_balance | units_of_production | none), a useful life in months, a residual value (basis points or absolute Rappen), the three GL accounts (cost=asset, accumulated depreciation=asset/liability contra, depreciation expense=expense), and an optional default cost centre. code is 1-20 chars, unique per workspace case-insensitively (duplicate_code). Wrong account types are refused with invalid_account_type before any write.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | ||
| name | Yes | ||
| description | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| glAssetAccountId | Yes | ||
| residualValuePct | No | ||
| usefulLifeMonths | No | ||
| depreciationMethod | No | ||
| defaultCostCenterId | No | ||
| residualValueRappen | No | ||
| glAccumDeprAccountId | Yes | ||
| glDeprExpenseAccountId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the tool creates a category with inherited defaults, and provides specific validation behaviors: code must be 1-20 chars and unique per workspace case-insensitively (duplicate_code), and wrong account types are refused with invalid_account_type before any write. This adds substantial transparency, though it does not mention idempotency or return values.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded with the core purpose. It packs a significant amount of information into three sentences without fluff. The first sentence is long and technically detailed, but every clause adds value, and the two follow-up sentences nicely summarize validation rules.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 13 parameters, no output schema, and no annotations, the description covers essential business logic, defaults, validation, and account rules. It leaves gaps around the interaction of residualValuePct vs residualValueRappen and does not explain return values, but overall provides enough context for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains the meaning of depreciationMethod (with enum values), usefulLifeMonths, residualValue (basis points or Rappen), the roles of the three GL accounts, and the optional defaultCostCenterId. It also specifies code length and uniqueness. However, it omits mention of the required name, workspaceId, and idempotencyKey, which are common but still undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Create') and resource ('fixed-asset category'), and goes further to explain what the category contains (defaults inherited by assets). This clearly distinguishes it from sibling tools like asset_category_update, asset_category_archive, and asset_category_list without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for creating a new category but does not explicitly contrast it with asset_category_update or state when not to use it. The context is clear, but no exclusions or alternatives are mentioned, so guidance is only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_category_getARead-only
Read one fixed-asset category by id, archived or not.
| Name | Required | Description | Default |
|---|---|---|---|
| categoryId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint=true already covers safety. The description adds one useful behavioral detail—'archived or not'—indicating the tool returns categories regardless of archive status. However, it omits error behavior, permission requirements, or response shape, so the added transparency is limited.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that front-loads the verb and resource. It contains no filler, repetition, or superfluous details, making it efficient and immediately scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with two required parameters and a read-only annotation, the description is adequate but not complete. It does not mention what the response contains, what happens if the category is not found, or how to locate the ID, which an agent might need given the large sibling tool set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate, but it only says 'by id,' which loosely maps to categoryId. It does not explain workspaceId, the relationship between the two parameters, or any format/constraints. This adds minimal meaning beyond the schema field names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read'), a specific resource ('fixed-asset category'), and a scoping constraint ('by id, archived or not'). This clearly distinguishes it from sibling tools like asset_category_list (listing) and asset_category_archive (mutation), making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case—fetching a single category when you have its ID—but does not explicitly mention alternatives or exclusions. An agent must infer that asset_category_list is for multiple categories and that create/update/archive are for mutations, as no routing is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_category_listARead-only
List the workspace fixed-asset categories, ordered by code. Optional active filter (true = only live, false = only archived) and a case-insensitive search over code and name. Accepts a savedViewId (G00 saved-view seam).
| Name | Required | Description | Default |
|---|---|---|---|
| active | No | ||
| search | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, covering the safety profile. The description adds behavioral details: ordering by code, case-insensitive search over code and name, and acceptance of savedViewId (G00 saved-view seam). However, it does not disclose default behavior (e.g., whether archived items are included by default) or any pagination/limits, which are typical for list operations. With readOnlyHint covering mutation concerns, a 3 is appropriate—adds some value but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no waste. The main action is front-loaded, and each clause adds a distinct piece of information (ordering, filters, search scope, savedViewId). The description is compact and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema, the description covers the essential behavior: what is listed, ordering, filters, and the savedView seam. It lacks explicit mention of default filter behavior and pagination, but these are common patterns and the readOnlyHint already signals safety. Overall, the description is adequate for an agent to call this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden. It explicitly explains all three optional parameters: 'active' (true = only live, false = only archived), 'search' (case-insensitive over code and name), and 'savedViewId' (G00 saved-view seam). The required workspaceId is self-evident from the context. This fully compensates for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists workspace fixed-asset categories, ordered by code. It distinguishes itself from sibling tools like asset_category_create, asset_category_get, and asset_category_archive by specifying a list operation with ordering. The verb 'List' and resource 'fixed-asset categories' are explicit and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: whenever a list of fixed-asset categories is needed, optionally filtered by active status or search. It does not explicitly name alternatives or when-not conditions, but the optional filters and mention of the savedViewId seam suggest clear use cases. The absence of explicit exclusions is a minor gap, but the context is sufficient for an agent to decide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_category_resolve_defaultsARead-only
Resolve the FULL default set for a category so an agent can create an asset (H01) with no further questions: the depreciation trio, the three GL accounts as id+number+name, and the optional cost centre. An archived category is refused with category_archived (choose another or un-archive first).
| Name | Required | Description | Default |
|---|---|---|---|
| categoryId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true, the bar is lower. The description discloses what data is returned (depreciation trio, GL accounts, cost centre) and the refusal behavior for archived categories (category_archived). This adds useful context beyond the annotation and is consistent with it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words. The primary purpose is front-loaded, and the error case is appended succinctly. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Since there is no output schema, the description properly explains the return content (the specific data fields) and the error condition. It asserts completeness ('FULL default set') and the goal of enabling asset creation without further questions, so nothing critical is missing for an agent to use this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions 'for a category' which implies categoryId, but does not explicitly define workspaceId or provide parameter-level detail. The two parameters are obvious from the tool name and context, but the description does not add much beyond what the schema names convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: 'Resolve the FULL default set for a category' and enumerates exactly what is resolved (depreciation trio, GL accounts, cost centre). It also clarifies the purpose ('so an agent can create an asset (H01) with no further questions'), which clearly distinguishes it from asset_category_create/update/archive siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use: when creating an asset, call this to obtain defaults. It also provides guidance on the archived-category error ('choose another or un-archive first'), which is a useful fallback. However, it does not explicitly mention alternatives or exclusions, though the use case is clear from the context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_category_updateA
Edit a fixed-asset category through a patch object. The same validation as create applies (account types, useful life, residual bounds); code is immutable through this verb. Only the fields present in patch change.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | No | ||
| categoryId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses that validation is the same as create (account types, useful life, residual bounds), that code is immutable, and that only fields present in the patch change. This is meaningful behavioral context beyond the schema. It doesn't mention idempotency behavior or error cases, but the disclosed traits are significant.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each carrying distinct information: what the tool does, validation parity with create, and the patch semantics. No filler, no repetition of schema field names. The most important scoping information (patch object, immutable code) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a patch-update tool with no output schema, the description covers the key behavioral aspects: what changes, what doesn't, and what validation applies. It doesn't mention idempotencyKey usage or response format, but those are less critical for an update operation. The description is complete enough for an agent to invoke correctly, though it could note that idempotencyKey is required for retries.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the 'patch' object semantics (only fields present change) and mentions validation categories, but it doesn't describe individual parameters like workspaceId, categoryId, or idempotencyKey. The patch object's fields are self-explanatory from the schema, but the required parameters' roles are not explained. Baseline 3 is appropriate because the description adds some meaning (patch semantics) but doesn't fully compensate for the 0% coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Edit'), a specific resource ('fixed-asset category'), and the mechanism ('through a patch object'). It also distinguishes itself from create by noting 'code is immutable through this verb' and that only fields present in patch change. This clearly differentiates it from sibling tools like asset_category_create and asset_category_archive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when you need to edit an existing fixed-asset category, and it explicitly notes that code cannot be changed here (implying create is the place for that). It also states that only fields present in the patch change, which is a key usage guideline. However, it doesn't explicitly name alternatives or state when NOT to use it (e.g., use create for new categories, use archive for deletion).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_createA
Create a capitalised fixed asset FROM a category (H00), inheriting its depreciation method, useful life, residual rule and three GL accounts; only name, acquisitionDate and acquisitionCostRappen (Rappen, >0) are required. Any of depreciationMethod/usefulLifeMonths/residualValuePct/residualValueRappen/the three GL account ids may OVERRIDE the inherited default at creation. number is optional (a unique FA-#### is generated if omitted; a manual one is unique per workspace case-insensitively, else duplicate_number). An archived category is refused with category_archived. The asset is created in status=draft; its financial baseline is immutable once a posted acquisition (H02) moves it out of draft.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| model | No | ||
| notes | No | ||
| number | No | ||
| barcode | No | ||
| categoryId | Yes | ||
| locationId | No | ||
| description | No | ||
| workspaceId | Yes | ||
| manufacturer | No | ||
| serialNumber | No | ||
| warrantyUntil | No | ||
| idempotencyKey | Yes | ||
| acquisitionDate | Yes | ||
| decliningRateBp | No | ||
| glAssetAccountId | No | ||
| residualValuePct | No | ||
| usefulLifeMonths | No | ||
| responsibleUserId | No | ||
| depreciationMethod | No | ||
| residualValueRappen | No | ||
| totalEstimatedUnits | No | ||
| glAccumDeprAccountId | No | ||
| acquisitionCostRappen | Yes | ||
| glDeprExpenseAccountId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so excellently. It states the draft status, immutability of the financial baseline after a posted acquisition, refusal of archived categories, number generation/uniqueness rules, and override behavior—far beyond a simple 'create' statement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, each earning its place: purpose/inheritance, required fields, overrides, numbering, error behavior, and lifecycle. Information is front-loaded and there is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 25-parameter creation tool with no output schema, the description is unusually complete: it covers required fields, defaults, overrides, error conditions, and post-creation lifecycle. It does not mention the return payload or explain idempotencyKey, which would have made it fully complete for an agent invoking a create operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description compensates by explaining the required fields, the Rappen unit and >0 constraint, the override list, and the optional number semantics. It leaves many optional parameters (e.g., locationId, serialNumber, idempotencyKey) undocumented, but the core creation-critical parameters receive meaningful semantic detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Create a capitalised fixed asset FROM a category (H00)', which clearly distinguishes it from related asset tools like asset_acquire or asset_update. It also explains the inheritance behavior that makes this tool unique among siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: creation requires a category and only name, acquisitionDate and acquisitionCostRappen. It does not explicitly name alternatives or state when not to use it, but the category-based inheritance semantics strongly imply its intended use case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_depreciation_forecastARead-only
A multi-period depreciation forecast (H09, US-H09.2) re-using the SAME pure H03 calculators the preview/schedule verbs use, so a forecast line never disagrees with a live asset_depreciation_schedule. fromPeriod/toPeriod are YYYY-MM inclusive; the window is at most 60 months and a reversed or over-long window is invalid_period_range. groupBy is none (default) | category | location (cost_center is not an asset-master dimension in Phase 1: invalid_input). filter narrows by categoryIds/locationIds/assetIds (a foreign assetId is not_found before any projection). Returns an ordered periods list, each with totalAmountRappen and an optional byGroup breakdown, a totalProjectedRappen, assetsReachingResidual, and (with includeDetail:true) per-asset ScheduleLine detail. A units_of_production asset with no unitsForecast contributes zero and raises a production_data_required warning rather than inventing usage. Writes nothing; deterministic for a given snapshot.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No | ||
| groupBy | No | ||
| toPeriod | Yes | ||
| fromPeriod | Yes | ||
| workspaceId | Yes | ||
| includeDetail | No | ||
| unitsForecast | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint=true annotation, the description adds 'Writes nothing; deterministic for a given snapshot' and details error/warning behaviors (invalid_period_range, invalid_input, not_found, production_data_required). This goes well beyond the annotation's read-only hint and informs an agent of failure and warning conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single dense paragraph carries no filler; every clause maps to a schema field, output field, or behavioral constraint. It front-loads the core purpose and consistency guarantee before diving into parameter semantics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies a detailed return contract: ordered periods, totalAmountRappen, optional byGroup, totalProjectedRappen, assetsReachingResidual, and includeDetail per-asset detail. All 7 parameters and error modes are addressed, making the description sufficient for accurate invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does comprehensively: date format (YYYY-MM), month-window cap, reversed/over-long error code, allowed groupBy values, filter dimensions, includeDetail semantics, and unitsForecast relevance. This gives the agent far more than the schema names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'multi-period depreciation forecast' with identifiers H09, US-H09.2. It also differentiates itself from sibling asset_depreciation_preview and asset_depreciation_schedule by describing the multi-period window and the guarantee of consistency with the live schedule.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context by explicitly referencing the 'preview/schedule verbs' and positioning this as the multi-period forecast that reuses the same pure calculators. It does not explicitly enumerate when-not-to-use or name the alternative to use instead, so it falls just short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_depreciation_methodsARead-only
List the registered depreciation methods (straight_line | declining_balance | units_of_production | none) with this workspace enablement flags. Each descriptor carries labelKey, descriptionKey, requiresUnits, requiresRate and enabled. A disabled method is hidden from the H00 category / H01 asset method pickers. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses meaningful behavior: output descriptor fields (labelKey, descriptionKey, requiresUnits, requiresRate, enabled, labelKey) and the effect of disabled methods on H00/H01 pickers. This meaningfully exceeds what the annotation already provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each contributing useful information: the resource, the method values, the descriptor fields, and the behavioral consequence of disabled flags. It is dense but not bloated, and the action is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by naming all returned descriptor fields and the allowed method enum values. It also clarifies the read-only nature and the visibility behavior, which is enough for an agent to understand what it will receive and how to interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required parameter workspaceId is not explicitly documented; the description only indirectly refers to 'this workspace.' With 0% schema description coverage, some compensation is expected, but workspaceId is a simple required string and the schema property name is mostly self-explanatory. The description does not add much parameter-level detail, so 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('List') and a clear resource: registered depreciation methods, and even enumerates the possible values: straight_line, declining_balance, units_of_production, none. It also identifies the returned descriptor fields, making it easy to distinguish this read-only listing tool from the sibling asset_depreciation_method_set_enabled.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives useful context: each descriptor carries enabled flags and disabled methods are hidden from pickers, so an agent can infer this is the appropriate tool for inspecting available depreciation methods. It does not explicitly name alternatives like asset_depreciation_method_set_enabled or list when not to use it, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_depreciation_method_set_enabledA
Enable or disable a depreciation method for this workspace (a setup action). Absence means enabled, so this only records a method switched off. The none method can never be disabled (a non-depreciating asset must always be expressible), refused with method_locked. An unknown methodKey is unknown_method. Idempotent on the (workspace, method) row.
| Name | Required | Description | Default |
|---|---|---|---|
| enabled | Yes | ||
| methodKey | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: it explains the absence-means-enabled semantics, that the operation only records a method switched off, that none can never be disabled and returns method_locked, that unknown methodKey returns unknown_method, and that the operation is idempotent. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, each contributing distinct information, with the primary action front-loaded and edge cases packed efficiently afterward. There is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no annotations, no schema parameter descriptions, and no output schema, the description covers the purpose, scope, edge cases, error conditions, and idempotency behavior. This is enough for an agent to understand when and how to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema supplies only names and types, so the description compensates well by clarifying the enabled semantics, the methodKey edge cases, the workspace scope, and idempotency behavior. It does not fully specify every parameter detail, such as the exact idempotencyKey format, so a small gap remains.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action: enabling or disabling a depreciation method for a workspace, and frames it as a setup action. It also clarifies the key semantic that absence means enabled, so the tool can be distinguished from the many asset and valuation method tools in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear operational context: it is a setup action scoped to a workspace and is the toggle for depreciation methods. It does not explicitly name alternatives such as asset_depreciation_methods or the parallel inventory_valuation_method_set_enabled, so it stops short of explicit alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_depreciation_previewARead-only
Preview the depreciation amount for period (YYYY-MM) for a list of assetIds, or for every active asset in the workspace when assetIds is omitted. Returns one DepreciationResult per asset: amountRappen (integer Rappen, half-away-from-zero), isFinal (this period lands NBV exactly on residual), projectedNbvAfterRappen (never below residual), and a reason when zero (already_at_residual | period_already_processed | non_depreciable | missing_production_data | unknown_method). Writes NOTHING and creates no journal (H04 posts). proRata is full_period (default) or actual_days; unitsByAsset maps assetId to the units produced this period (units_of_production); paramsByAsset carries per-asset decliningRateBp / totalEstimatedUnits overrides. A foreign or unknown assetId is rejected with not_found before any calculation (§H-TENANT).
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| proRata | No | ||
| assetIds | No | ||
| daysByAsset | No | ||
| workspaceId | Yes | ||
| unitsByAsset | No | ||
| paramsByAsset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses significant behavioral details: it returns a reason when the amount is zero with enumerated possible reasons, states that projectedNbvAfterRappen never goes below residual, describes the rejection of foreign/unknown assetIds with not_found, and explains the effect of proRata and override parameters. This adds rich context beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-organized: it starts with the core purpose, then return structure, side-effect guarantee, parameter details, and error handling. Every sentence adds value, and there is no fluff. It is a single paragraph but logically segmented, making it effective despite its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of 7 parameters, nested objects, and no output schema, the description covers return types, parameter semantics, error cases, and side-effect behavior. It provides enough detail for an agent to call the tool correctly without additional documentation, including the reason enums and the guarantee about NBV never going below residual.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by explaining most parameters: period format (YYYY-MM), assetIds (omission means all active), proRata (full_period or actual_days), unitsByAsset (mapping to units produced), and paramsByAsset (overrides for decliningRateBp and totalEstimatedUnits). However, daysByAsset is not explicitly described, and the exact structure of nested objects could be clearer, so it is not fully complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: preview depreciation amount for a period, either for specific assetIds or all active assets. It specifies the verb 'preview' and the resource (depreciation amount), and distinguishes itself from posting tools by explicitly noting it writes nothing and creates no journal, which separates it from siblings like asset_depreciation_run_create.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly guides usage by stating it is a read-only preview that creates no journal, implying it should be used when the agent wants to estimate depreciation without committing. It also explains the behavior when assetIds is omitted (covers all active assets). However, it does not explicitly name alternative tools for posting or other scenarios, so guidance is not fully explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_depreciation_run_createA
Create a DRAFT depreciation run for a period (H04): calculate the depreciation for every eligible asset via the H03 engine and persist a run header (status=draft) plus one line per asset that returned amount > 0, so a bookkeeper can review the proposed expense before anything hits the books. Eligible = status active/fully_depreciated, method != none, accumulated depreciation below cost minus residual, last_depreciation_period null or before the target period, not disposed. Optional filters narrow the set: assetIds (an explicit subset), categoryId, costCenterId. unitsByAsset maps assetId to the WHOLE units produced this period and is what makes a units_of_production asset depreciable at all (TILL captures no production data of its own, so the run carries it; same shape as asset_depreciation_preview, and values must be non-negative safe integers up to 1e12 or the create is invalid_input). skipped[] reports assets that produced NO line, as {assetId, assetNumber, reason}: on an EXPLICIT assetIds list every named asset that produced no line is reported (non_depreciable | asset_terminal | already_at_residual | period_already_processed | filtered_out | missing_production_data | missing_units_estimate | zero_amount), because naming an asset is an instruction about that asset; on a SWEEP the eligible register IS the selection, so only assets that reached the calculation are reported (missing_production_data | missing_units_estimate | zero_amount). A units_of_production asset named EXPLICITLY in assetIds is refused, not merely reported, when the input can be corrected: missing_production_data (no figure supplied) or missing_units_estimate (the asset carries no totalEstimatedUnits, so there is no denominator to allocate over). Nothing is written in either case. postingGranularity is detailed (one expense + one accum line per asset, the default) or summarised (grouped by account + cost centre); the per-asset lines exist either way for the sub-ledger audit. period is YYYY-MM. Refused with invalid_period, period_locked (the period is hard-locked in A03) or run_already_exists (a draft for the same period and selection already exists, and the response names it). Writes NO journal (the post verb does). Idempotent on idempotencyKey and on the (period, selection) signature, where the signature is built from the LINES the run would write (asset, amount and production figure), so two creates over the same eligible set yield ONE draft however the filters or the units map were spelled. When there is nothing to charge, NO run is persisted at all: the answer is {ok:true, run:null, lines:[], empty:true, skipped:[...]}, computed fresh, so it is identical on every call however many times it is repeated and whatever key each call carries. A run header exists only where at least one asset is charged, which is why a finished period never leaves an outstanding draft in run_list for a period-close checklist to trip over. Once a period is POSTED its assets carry it as their last depreciation period, so nothing is eligible and a later create for that period returns exactly that empty answer, carrying alreadyPostedRunId: the id of the most recently posted run for the period (two disjoint selections can both post for one period). A foreign assetId is not_found (§H-TENANT).
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| assetIds | No | ||
| categoryId | No | ||
| workspaceId | Yes | ||
| costCenterId | No | ||
| unitsByAsset | No | ||
| idempotencyKey | Yes | ||
| postingGranularity | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With zero annotations, the description carries the full burden and delivers richly: it discloses side effects (persists header and lines), non-effects ('Writes NO journal', 'Nothing is written in either case'), idempotency on both idempotencyKey and the (period, selection) signature, and the no-op empty outcome where 'NO run is persisted at all'. It even discloses state-transition behavior ('Once a period is POSTED its assets carry it as their last depreciation period') and error semantics (period_locked, run_already_exists, not_found §H-TENANT).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence carries real operational meaning and the length is arguably justified by the tool's complexity (8 params, two filter modes, two granularities, many error codes). However, it is a single dense unbroken paragraph with no paragraph breaks or bullets, and there is some redundancy (the 'identical on every call' restatement of idempotency and the period-close checklist rationale). The size is appropriate; the structure is not.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 8 parameters, no annotations, and no output schema, the description covers eligibility, filters, skipped-reporting semantics, granularity, error codes, idempotency, and empty-run behavior in exceptional depth. The remaining gaps are the explicit success-case return shape (only the empty response {ok:true, run:null, lines:[], empty:true, skipped:[...]} is fully spelled out) and the thin treatment of workspaceId's scoping. These are minor but real contract gaps given the absence of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does comprehensively: period is 'YYYY-MM'; unitsByAsset is defined as whole units produced with 'non-negative safe integers up to 1e12' and validation consequences; postingGranularity is explained as detailed vs summarised with sub-ledger implications; idempotencyKey's role is described; and assetIds/categoryId/costCenterId are identified as narrowing filters with distinct skipped-reporting semantics. Only workspaceId's tenant-scoping is left to the §H-TENANT cross-reference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Create a DRAFT depreciation run for a period (H04)') and states exactly what is persisted: a run header with status=draft plus one line per asset with amount > 0. It also distinguishes itself from its closest siblings by noting 'Writes NO journal (the post verb does)' and referencing asset_depreciation_preview for the unitsByAsset shape. An agent can unambiguously separate this from run_post, run_reverse, run_get, and preview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The draft-vs-post lifecycle is explicit ('so a bookkeeper can review the proposed expense before anything hits the books'; 'the post verb does'), which tells the agent when to call this tool versus asset_depreciation_run_post, and the full eligibility criteria define when a create is meaningful. However, there is no explicit 'use X instead when you don't want to persist' statement; the preview alternative is only implied via the shape reference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_depreciation_run_getARead-only
Read one depreciation run by id (H04): the run header (period, status, granularity, total, asset count, the journal_entry_id it posted and the reversing link if reversed) plus every line, in register order (asset id, assetNumber, amount, accumulated-before/after, nbv-after, is_final, costCenterId, the two GL account ids the line posts against, and unitsProduced: the production figure a units_of_production amount was computed from, null for every other method). §H-TENANT: a foreign runId is not_found.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint=true already marks this as read-only, and the description reinforces this with 'Read'. More importantly, it goes beyond the annotation by detailing the exact data payload (header fields, line fields, ordering) and disclosing a specific error condition (foreign runId returns not_found). This is valuable behavioral transparency that helps the agent set expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but well-structured: it leads with the primary purpose, then itemizes header fields, then line fields, and ends with a tenant-specific note. Every sentence adds information, and the format is easy to scan. It is not overly verbose given the amount of detail needed to enumerate the response.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without an output schema, the description must stand alone in explaining what the tool returns. It does so exhaustively, listing all header and line attributes including nested items like GL account ids and unitsProduced. It also covers edge cases (reversing link, foreign runId). Given the tool's complexity, the description is remarkably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the burden. It implicitly links runId to the id being read and mentions 'foreign runId' which clarifies scope. The workspaceId is not elaborated, but the tenant note implies it's the tenant context. It does not explicitly map each parameter, but it provides enough to understand their roles in this context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Read'), a clear resource ('one depreciation run by id'), and an explicit objective. It enumerates exactly what will be returned (header fields plus every line with all fields), leaving no ambiguity about the tool's function. It is distinct from sibling tools like asset_depreciation_run_list or create because 'Read' implies a singular fetch by id.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies usage when the agent has a specific runId and needs the full detail of a single run. It does not explicitly name alternatives or state when not to use it, but the verb and resource make the intent unambiguous. Given the large sibling set, an explicit cross-reference to list or create would elevate it, but the current text suffices.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_depreciation_run_listARead-only
List the workspace depreciation runs (H04), newest period first. Optional filters: period (exact YYYY-MM), status (draft | posted | reversed, one or an array), and a from/to period window. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| period | No | ||
| status | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already marks the tool safe, and the description reinforces it with 'Read-only.' Beyond the annotation, it adds behavioral specifics: ordering newest first, exact period format, and allowed status values. It does not mention pagination or response shape, but the added context meaningfully exceeds annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with no filler. The primary purpose is front-loaded, followed by the optional filters and a closing read-only note. Every sentence contributes information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential selection context: what is listed, ordering, filters, and status vocabulary. Since there is no output schema, the agent still lacks guidance on returned fields, pagination, or response structure, which is a meaningful gap for a list endpoint.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry parameter meaning. It does explain period format, status values, and the from/to window. However, it leaves workspaceId implicit and does not define the exact format or inclusivity of the from/to endpoints, so coverage is useful but incomplete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), a precise resource ('workspace depreciation runs (H04)'), and an ordering guarantee ('newest period first'). This clearly distinguishes it from sibling tools like asset_depreciation_run_get or asset_depreciation_run_create by establishing that this is the read-only list operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description communicates clear context: list runs for a workspace, optionally filtered by period, status, and date window. However, it does not explicitly say when to prefer this over closely related siblings such as asset_depreciation_run_get (single run retrieval) or asset_depreciation_preview/schedule (projections), leaving the selection partly implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_depreciation_run_postA
Post a reviewed DRAFT depreciation run (H04): in ONE atomic transaction it re-validates eligibility and the period-open state, posts one balanced GL journal via A02 (Dr the depreciation-expense account / Cr the accumulated-depreciation account for each line, source=asset_depreciation), writes one append-only asset_transaction (type=depreciation) per asset, advances every asset accumulated_depr_rappen / net_book_value_rappen / last_depreciation_period, sets a final asset to status=fully_depreciated, and marks the run posted with its journal_entry_id. postingDate defaults to the last day of the period; it may land later than the period (a late close) but never earlier than the first day of the run period, and an explicitly supplied date may not predate the acquisition date of any asset in the run (invalid_input names the field and a reason of before_period or before_acquisition). Refused with period_locked (the RUN period or the posting date is hard-locked: a posting date in an open year can never slip a charge past a sealed one), stale_draft with the reason that applies (accumulated_moved: its accumulated depreciation moved since the draft was calculated; valuation_moved: its net book value moved, for instance an asset_add_capitalisation landed while the draft was under review; asset_terminal: it was disposed or archived; asset_missing), run_reversed. A stale draft is never recomputed, because that would silently change a figure the reviewer already approved: recreate the run. Idempotent on idempotencyKey AND on run status: a re-post of an already-posted run posts NO second journal and writes NO second asset_transaction, returning the original result (a double-post never double-counts). §H-TENANT: a foreign runId is not_found. CONSEQUENCE: Posts the depreciation run across every asset in it; a double-post never double-counts.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| postingDate | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since no annotations are provided, the description carries the full burden of behavioral disclosure. It fully covers atomicity, the nested steps (journal posting, asset_transaction writes, ledger updates), idempotency (no double-counting), date constraints, and all failure modes with specific reasons (stale_draft, period_locked, run_reversed). This is exemplary transparency for a complex mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is lengthy but every sentence adds value. It front-loads the core purpose and then details behaviors and errors. It could be tightened slightly (e.g., merging some redundant phrases like 'a double-post never double-counts' repeated), but overall it is well-organized and efficiently communicates a high-complexity operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and sparse schema, this description is remarkably complete. It covers side effects, error conditions with reasons, idempotency, tenant scoping, and date semantics. An agent has all necessary information to call the tool correctly and interpret failures, which is exceptional given the operation's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema parameter descriptions, the description must compensate. It extensively explains postingDate (defaults, late-close rules, constraints vs acquisition date) and mentions idempotencyKey in the idempotency context. It does not explicitly elaborate on workspaceId or runId, but these are self-evident from the tool name and usage, so the coverage is mostly adequate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Post a reviewed DRAFT depreciation run (H04)' and enumerates the full scope of what the operation does (re-validates eligibility, posts a balanced GL journal, writes asset_transaction rows, updates balances, marks the run posted). It also differentiates from sibling tools like asset_depreciation_run_reverse and asset_depreciation_run_create by framing itself as the post action on a reviewed draft.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states the precondition (a reviewed DRAFT run) and enumerates refusal reasons and the instruction to 'recreate the run' for stale drafts. It does not explicitly name the create/reverse sibling tools as alternatives, but the context makes the intended use clear. The guidance is strong but not exhaustive about when to use this versus related tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_depreciation_run_reverseA
Reverse a POSTED depreciation run (H04) when a material error is found, without rewriting history: it calls A02 reverseEntry on the original journal (producing a new reversing entry), writes a compensating asset_transaction (type=depreciation_reversal) per line, restores every asset accumulated_depr_rappen / net_book_value_rappen / status / last_depreciation_period to its pre-run value (by delta), and marks the original run reversed with the reversing journal link. The original journal and the original run amounts are never mutated (§H-AUDIT). reverseDate defaults to the last day of the period and may not predate the first day of the run period (a reversal is booked forward, never before the charge it reverses: invalid_input names the field). The optional reason is recorded on the reversing entry description and on the compensating asset_transaction. Refused with run_not_posted (a draft), already_reversed, period_locked, and later_run_exists when a LATER run has already posted for one of the assets (the response names the blocking period): runs are reversed newest first, because the asset last_depreciation_period stays on the later run and the earlier period would otherwise be silently unrecoverable. A reversed run cannot be re-posted; a new run for the same period may be created afterwards. Idempotent on idempotencyKey. §H-TENANT: a foreign runId is not_found. CONSEQUENCE: Posts a reversing entry against a posted depreciation run, without rewriting history.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| reason | No | ||
| reverseDate | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and delivers richly: it discloses the A02 reverseEntry call, the compensating asset_transaction type, the exact asset fields restored, the audit constraint (§H-AUDIT), the default reverseDate behavior, the idempotency on idempotencyKey, tenant scoping (§H-TENANT), and the consequence. It even names the error cases and the blocking period behavior. This is far beyond what annotations would typically provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-rich, with the core action front-loaded. Every sentence adds meaningful behavioral or constraint information. It loses one point because it is long and somewhat sprawling, mixing audit references, error cases, and consequences in a single paragraph; a structured breakdown would improve scannability. However, no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex mutation tool with no annotations and no output schema, the description covers: what happens (reversing entry, compensating transaction, asset restoration), the constraints (date, ordering, idempotency, tenant), the refusal conditions, and the post-condition (cannot re-post, new run allowed). An agent has everything needed to decide when to call it and what to expect. The only minor omission is the return value shape, but with no output schema the description still provides sufficient behavioral context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does: reverseDate semantics are fully explained (defaults to last day of period, may not predate first day), reason is described as recorded on the reversing entry and compensating transaction, and idempotencyKey is explained as making the operation idempotent. runId and workspaceId are not individually detailed, but their roles are clear from the context (foreign runId is not_found, workspaceId scopes the tenant). Minor gap: no explicit format for reverseDate, but the semantic guidance is strong.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Reverse a POSTED depreciation run (H04)') and names the exact resource and context. It distinguishes itself from sibling tools like asset_depreciation_run_post and asset_depreciation_run_create by stating it reverses a posted run, and it explicitly contrasts with the generic reverse_entry sibling by describing the compensating asset_transaction and asset state restoration. The scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use it ('when a material error is found'), and provides clear exclusions: refused with run_not_posted, already_reversed, period_locked, and later_run_exists. It also states the ordering constraint ('runs are reversed newest first') and that a reversed run cannot be re-posted. This is explicit when/when-not guidance with named alternatives implied by the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_depreciation_scheduleARead-only
Project the full remaining depreciation schedule for ONE asset from fromPeriod (YYYY-MM, inclusive) to an optional toPeriod, using the same pure engine as the preview so a line never disagrees with a live preview. Returns ordered ScheduleLine rows (period, amountRappen, projectedAccumRappen, projectedNbvRappen, isFinal); the final line residual-adjusts so the lines sum to exactly cost minus residual. A units_of_production asset needs a unitsForecast map (period to units) or the projection returns incomplete with a units_forecast_required warning rather than inventing usage. Writes nothing; a foreign assetId is not_found (§H-TENANT).
| Name | Required | Description | Default |
|---|---|---|---|
| params | No | ||
| assetId | Yes | ||
| proRata | No | ||
| toPeriod | No | ||
| fromPeriod | Yes | ||
| workspaceId | Yes | ||
| unitsForecast | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description explicitly says 'Writes nothing' and details edge-case behavior: the final line residual-adjusts to match cost minus residual, units_of_production assets require a unitsForecast or return an incomplete projection with a warning, and a foreign assetId yields a not_found error. This is rich behavioral disclosure that far exceeds the annotation's minimal signal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense yet efficient, with every sentence contributing distinct value: purpose, output shape, residual behavior, units_of_production caveat, and safety/error semantics. It is front-loaded with the core action and remains readable without unnecessary jargon.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema and 0% param coverage, the description provides a complete-enough contract: return row fields, residual adjustment guarantee, unitsForecast requirement, and error code reference. It leaves little ambiguous for an agent deciding whether and how to call this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates well by explaining fromPeriod's format and inclusivity, toPeriod as optional, and unitsForecast's required shape for units_of_production assets. It does not explain proRata or params, but the most critical parameters are semantically defined in plain language, which is a strong addition given the schema is silent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Project the full remaining depreciation schedule for ONE asset', and immediately distinguishes itself from the sibling preview tool by noting it uses 'the same pure engine as the preview so a line never disagrees with a live preview'. This makes the tool's purpose unmistakable and separates it from asset_depreciation_forecast or asset_depreciation_preview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: it projects the full remaining schedule for one asset over a date range, and explicitly ties its behavior to the preview engine for consistency. It does not name alternatives or state when not to use it, but the scoping to 'ONE asset' and the optional toPeriod strongly imply when this tool is appropriate vs. other asset tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_disposal_getARead-only
Read one fixed-asset DISPOSAL sub-ledger transaction by id (H06), with the balanced GL journal entry it posted. transactionId must name a disposal-type asset_transaction in this workspace: a foreign id, or an id naming an acquisition / depreciation event, is not_found rather than a leak of the wrong event through the disposal verb (§H-TENANT).
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| transactionId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers non-mutation. The description adds meaningful behavior beyond that: not_found on invalid ids, cross-tenant concern (§H-TENANT), and the fact that the returned data includes the posted GL journal entry.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with no filler. The purpose is front-loaded, then the error/scope constraint is added precisely. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple get-by-id with no output schema, the description states the return contains the posted GL journal entry and explains invalid-id behavior. It does not explicitly walk through workspaceId, but the tool is simple enough that the description is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the burden and mostly succeeds: it explains transactionId must name a disposal-type asset_transaction and what happens with foreign or wrong-typed ids. workspaceId is only implied as the workspace scope via 'in this workspace', which is useful but less explicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Read'), a specific resource ('fixed-asset DISPOSAL sub-ledger transaction by id (H06)'), and the key return content ('balanced GL journal entry'). The 'DISPOSAL' qualifier clearly differentiates it from generic asset_transaction_get and from disposal preview/dispose verbs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit exclusion semantics: transactionId must be a disposal-type asset_transaction, and foreign or acquisition/depreciation ids return not_found rather than leaking the wrong event. This tells an agent when not to use the tool, though it does not name a specific alternative sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_disposal_previewARead-only
Preview the exact disposal journal for an active or fully_depreciated fixed asset (H06) WITHOUT posting anything. Given the asset, a disposalDate, proceedsRappen (>= 0), a gainLossAccountId (an income or expense account) and, when proceeds > 0, a proceedsAccountId (a bank/receivable asset or a liability account), it returns the current cost / accumulated depreciation / net book value, the signed gainLossRappen (proceeds minus NBV: positive is a gain, negative a loss), and the full balanced journal_lines it WOULD post: Dr accumulated depreciation, Dr proceeds account, the book loss (Dr) or gain (Cr) on the gainLossAccountId, and Cr the asset cost account. The lines are the SAME lines asset_dispose posts, so the preview never disagrees with the posting. Pure read: it writes nothing, creates no journal and does NOT check the period lock, so it is exactly the tool for seeing what a disposal would book before finding out the period is sealed. Refused before any calculation with not_found (foreign asset, §H-TENANT), asset_already_disposed / asset_archived / asset_not_acquired (a draft has no cost basis), invalid_proceeds (negative or non-integer), invalid_gain_loss_account or invalid_proceeds_account (missing, foreign, or wrong type).
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| reason | No | ||
| assetId | Yes | ||
| workspaceId | Yes | ||
| disposalDate | Yes | ||
| proceedsRappen | Yes | ||
| counterpartyName | No | ||
| gainLossAccountId | Yes | ||
| proceedsAccountId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true, so the description's emphasis on 'WITHOUT posting anything' and 'Pure read' aligns with that. It goes beyond annotations by detailing what the tool does NOT do (does not check period lock) and what it returns (detailed journal lines). This adds context about execution behavior and side-effect absence, which is valuable. However, it could have also disclosed potential performance or rate-limit traits, but that's not critical.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured, front-loading the core purpose and key constraints, then detailing the return value and refusal reasons. Each sentence adds critical information with no fluff. It's longer than average, but every clause earns its place given the complexity of the disposal logic.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of asset disposal (multiple accounts, conditional parameters, validation rules), the description is comprehensive. It explains what the tool returns, how the journal lines are constructed, and all refusal scenarios. The only minor gap is that it doesn't explicitly describe the output schema for the journal_lines structure (e.g., exact field names), but since there is no output schema defined and the description already provides a textual breakdown, it is sufficiently complete for an agent to understand the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% in structured form, so the description compensates by explaining each key parameter's role, especially the nuanced validation (proceedsRappen >= 0, gainLossAccountId type, proceedsAccountId conditional). It clarifies the purpose of optional fields like notes/reason/counterpartyName implicitly by listing them in the schema, but the description adds meaningful context for the core parameters. It does not detail every parameter but covers the ones essential for correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the tool's purpose: to preview the disposal journal for an asset without posting. It specifies the exact resource (fixed asset H06), the action (preview), and distinguishes it from the posting tool. It names the key output elements (lines, gain/loss calculation) and explains the 'same lines' guarantee, which is clear and non-ambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when to use this tool (to see what a disposal would book before posting, especially if period lock is unknown) and contrasts it with asset_dispose. It also provides explicit conditions for required parameters (e.g., proceedsAccountId needed when proceeds > 0) and lists all refusal reasons, giving agents complete routing information.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_disposal_summaryARead-only
The disposal gain/loss pack for a date range (H09, US-H09.4): every H06 disposal posted between fromDate and toDate (ISO YYYY-MM-DD, inclusive) with its asset number/name, disposalDate, originalCostRappen, accumDeprAtDisposalRappen, nbvAtDisposalRappen, proceedsRappen, gainLossRappen (signed: proceeds minus book value, so positive is a gain), journalEntryId and reason (the disposal description). Footer totals: count, sum proceeds, sum gain, sum loss, net gain/loss. Optional filter by categoryIds, locationIds and a case-insensitive reason substring. No disposals in range is an empty success, not an error. Pure read over the append-only sub-ledger, §H-TENANT.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No | ||
| toDate | Yes | ||
| fromDate | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description adds valuable behavioral context: it is a 'pure read over the append-only sub-ledger', and it explicitly states that 'no disposals in range is an empty success, not an error'. It also explains the sign convention of gainLossRappen. This fully discloses behavior beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient. It front-loads the purpose, then lists output fields, footer totals, and filters in a logical flow. Every sentence adds necessary information without fluff. It is appropriately sized for the complexity of the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, zero schema description coverage, no output schema, and a large sibling list, the description covers all essential aspects: input parameters, output fields, footer totals, filtering options, empty result behavior, and read-only nature. An agent can call it correctly and interpret results without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description fully compensates. It specifies date format (ISO YYYY-MM-DD) and inclusivity, explains the filter semantics (categoryIds, locationIds, case-insensitive reason substring), and clarifies the output fields. WorkspaceId is not described but is a common contextual parameter, likely understood from the environment.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (implied get), resource (disposal gain/loss pack), and scope (date range). It names the exact record set (H06 disposals) and lists output fields. This clearly distinguishes it from related tools like asset_disposal_get or asset_disposal_preview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description specifies the date range and optional filters, making the tool's use case clear. It does not explicitly name alternative tools or say when not to use it, but the context of a summary report over a range is evident from the wording. It lacks explicit exclusions but provides enough guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_disposeA
Dispose an active or fully_depreciated fixed asset (H06): the TERMINAL financial event (sale, scrap, donation, write-off, retirement). In ONE atomic transaction it (a) posts ONE balanced GL journal via A02 (source=asset_disposal) clearing the asset cost (Cr) and its accumulated depreciation (Dr), recognising the proceeds (Dr the proceedsAccountId) and the resulting book gain (Cr) or loss (Dr) on the gainLossAccountId; (b) writes ONE append-only asset_transaction (type=disposal, deltaCostRappen = -cost, deltaAccumDeprRappen = -accumulated, proceedsRappen, gainLossRappen) naming that journal; and (c) moves the asset to status=disposed, forces net_book_value_rappen to 0 and stamps disposed_at / disposal_proceeds_rappen (cost and accumulated stay for historical reporting). The gain/loss SIGN is proceedsRappen minus net book value: proceeds above NBV credit the gain account, below NBV debit the loss account, exactly equal writes no gain/loss line. A scrap (proceedsRappen = 0) writes no proceeds line and needs no proceedsAccountId. After success the asset is permanently excluded from every future depreciation run (H03/H04 status filter) and asset_update refuses it with asset_terminal. Refused with not_found (foreign asset, §H-TENANT), asset_already_disposed (a second dispose: first wins, second is refused), asset_archived, asset_not_acquired (a draft has no cost basis), invalid_proceeds, invalid_gain_loss_account, invalid_proceeds_account, or period_locked (disposalDate in a hard-locked A03 period). A wrong disposal is corrected by reversing the journal (A02) plus a compensating asset_transaction, never an un-dispose. Idempotent on idempotencyKey: a replay returns the original {asset, transaction, journalEntry, gainLossRappen} and posts no second journal, appends no second transaction and re-flips no status. CONSEQUENCE: Posts the asset's disposal and the resulting gain or loss; the asset leaves every future depreciation run.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| reason | No | ||
| assetId | Yes | ||
| workspaceId | Yes | ||
| disposalDate | Yes | ||
| idempotencyKey | Yes | ||
| proceedsRappen | Yes | ||
| counterpartyName | No | ||
| gainLossAccountId | Yes | ||
| proceedsAccountId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden and does so richly: it specifies the atomic single GL journal posting with exact debit/credit flows, the append-only asset_transaction fields, the status flip, forced net_book_value to 0, permanent exclusion from depreciation, asset_update refusal, all error codes, correction path, and idempotency semantics. This goes far beyond any annotation could.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but earned: it uses a clear (a)/(b)/(c) enumeration of the atomic steps, states gain/loss sign, scrap rules, consequences, error codes, correction, and idempotency in organized order. The opening and final 'CONSEQUENCE' sentence overlap slightly, but for a tool with this complexity every sentence carries operational value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must define the return value, which it does: idempotent replay returns {asset, transaction, journalEntry, gainLossRappen}. It also covers all side effects, failure modes, and post-conditions, making the tool fully callable by an agent without needing additional documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does for the financial core: proceedsRappen sign convention, proceedsAccountId absence for scrap, gainLossAccountId behavior, idempotencyKey replay, and disposalDate's role in period_locked. Notes, reason, and counterpartyName are not explained, but they are peripheral compared to the accounting semantics that the description does add.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Dispose an active or fully_depreciated fixed asset (H06): the TERMINAL financial event (sale, scrap, donation, write-off, retirement).' It distinguishes this from siblings like asset_disposal_preview and asset_disposal_get by emphasizing terminality, so an agent can tell exactly what it does and how it differs without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: it must be called on active or fully depreciated assets, is terminal, and correction is via reversal of the journal plus a compensating asset_transaction, never an un-dispose. It does not explicitly name alternative tools (e.g., asset_disposal_preview) for pre-checking, so no when-not-to-use guidance, but the context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_end_of_life_listARead-only
Assets at or approaching end of life (H09, US-H09.8): with status="approaching" (default) it projects each active, depreciable asset forward from the current period and returns those whose remaining life falls inside withinMonths (default 12, capped 60), each with current NBV, residual, remainingMonths and the next forecasted depreciation amount, ordered soonest first. With status="fully_depreciated" it lists the already-fully-depreciated assets. Re-uses the pure H03 projector; writes nothing. §H-TENANT.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| workspaceId | Yes | ||
| withinMonths | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, it explains that the tool projects active depreciable assets forward from the current period, filters by remaining life within a capped window, orders results soonest first, and writes nothing. This gives the agent a clear safety and behavior profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded and dense with useful facts, with no filler. It is a long single sentence with some domain shorthand (H09, §H-TENANT) that adds context but could be slightly tightened for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list/report tool with no output schema, it tells the caller exactly what comes back: current NBV, residual, remainingMonths, and next forecasted depreciation, plus ordering. The required parameter, status modes, and withinMonths cap are all covered, so no critical usage information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the parameter explanation: it defines status values and withinMonths default/cap. It does not explicitly describe workspaceId semantics, but that parameter is self-evident and named as required in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (assets at or approaching end of life), states the supporting report codes H09/US-H09.8, and contrasts the approaching vs fully-depreciated views. This is enough to set it apart from generic asset listing tools such as asset_list or asset_depreciation_schedule.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly explains the default mode, when to use status='fully_depreciated', and how withinMonths affects results. It gives clear context for using this tool, though it does not explicitly name sibling alternatives or exclusions, which keeps it from a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_getARead-only
Read one fixed asset by id, including its resolved financial baseline, the three GL accounts it will post against, status and the convenience columns (accumulated depreciation, net book value) later specs maintain.
| Name | Required | Description | Default |
|---|---|---|---|
| assetId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description's 'Read' matches. It adds value by disclosing the specific return contents (resolved financial baseline, GL accounts, convenience columns), which is more than a generic read. It does not mention auth or rate limits, but for a read-only tool with the annotation, the bar is met.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One efficient sentence that front-loads the purpose and then enumerates the return fields. No filler or redundancy. The trailing 'later specs maintain' is a minor artifact but does not detract.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with two parameters, no output schema, and a readOnlyHint annotation, the description adequately covers what the agent needs: it states the resource, the identifier, and the contents of the response. It is missing nothing critical for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies that the tool fetches 'by id', implicitly mapping to assetId, but does not explain workspaceId at all. The parameter names are self-explanatory, but the description adds minimal semantic value beyond the schema for these two parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and resource ('one fixed asset by id'), and lists the exact contents returned (financial baseline, three GL accounts, status, convenience columns). This clearly distinguishes it from siblings like asset_list (list all) and asset_search (filtered search) without needing to open their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: it is the tool for fetching a single fixed asset by its id. It does not explicitly mention alternatives or when not to use it, but the singular-by-id scope is unambiguous against the many asset-related siblings (asset_create, asset_update, asset_transaction_get, etc.).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_ledger_getARead-only
The complete chronological ledger for ONE fixed asset (H07, US-H07.1): every asset_transaction event (acquisition, additional_capitalisation, opening, depreciation, disposal) in effective-date order, each with its type, delta cost, delta accumulated depreciation, proceeds/gain-loss (disposal), the GL journal_entry_id it posted, and the RUNNING cost / accumulated depreciation / net book value AFTER that event. A disposed asset keeps its full history (its disposal row returns the running totals to zero; nothing is purged). Read-only. §H-TENANT: a foreign assetId is not_found.
| Name | Required | Description | Default |
|---|---|---|---|
| assetId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses meaningful behavior: disposed assets retain full history and are not purged, disposal rows reset running totals to zero, entries include the posted GL journal_entry_id, and a foreign assetId returns not_found. This is rich behavioral context not available anywhere else.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and front-loaded with the core purpose, then lists event types, output fields, running-total semantics, disposal behavior, read-only status, and tenant behavior. Each sentence adds substantive value; there is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only single-asset ledger tool with only two simple parameters and no output schema, the description provides enough detail to invoke it correctly: what events are included, ordering, exact fields, running totals, post-disposal behavior, and not_found semantics. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no parameter descriptions, so the description carries the burden. It clarifies that assetId identifies one fixed asset and adds tenant/error semantics ('foreign assetId is not_found'), but workspaceId remains unexplained beyond its self-evident name. Partial compensation for the 0% schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies a precise verb and resource: retrieving the complete chronological ledger for a single fixed asset. It enumerates the event types and fields returned, and the 'ONE fixed asset' scope clearly differentiates it from list-style siblings like asset_ledger_list or asset_transaction_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context through the output spec—use this when you need the full event ledger with running totals for one asset—but it does not explicitly state when to prefer it over alternatives such as asset_transaction_get or asset_get. No when-not-to-use guidance or alternative routing is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_ledger_listARead-only
List fixed-asset sub-ledger events ACROSS assets (H07, US-H07.4), newest first, with journal links. Optional filters: assetId, type (one or an array of acquisition | additional_capitalisation | opening | depreciation | disposal | revaluation | adjustment), a fromDate/toDate window (YYYY-MM-DD, inclusive), and journalEntryId. Simple limit (default 100, max 500) / offset pagination; returns { items, total }. This is a flat list, so no running balance is attached (the running NBV is only meaningful within one asset: use asset_ledger_get for that). Read-only, §H-TENANT.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | ||
| limit | No | ||
| offset | No | ||
| toDate | No | ||
| assetId | No | ||
| fromDate | No | ||
| workspaceId | Yes | ||
| journalEntryId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint=true already signals safety, and the description adds valuable behavioral context: it returns a flat list with no running balance, describes the response shape {items, total}, and mentions pagination defaults. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence contributes: scope, filters, pagination, return shape, a caution about running balance, and read-only flag. It could be trimmed slightly but remains efficient and front-loaded with the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description defines the return format. It covers all parameters, pagination behavior, and the key limitation (no running balance), pointing to the alternative. For an 8-parameter tool, nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description compensates by enumerating all optional filters (assetId, type with enum values, date window with format, journalEntryId) and pagination (limit default 100 max 500, offset). It even clarifies that type accepts a single value or an array. This fully bridges the gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (list), resource (fixed-asset sub-ledger events), scope (across assets), ordering (newest first), and inclusion of journal links. It explicitly distinguishes from asset_ledger_get by noting this is a flat list across assets, so an agent can tell them apart without inspecting schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit guidance on when to use this tool vs. the alternative: 'use asset_ledger_get for that' for running balance. It also states the scope (ACROSS assets) and the read-only nature, giving clear context for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_listARead-only
List the workspace fixed-asset register, ordered by number. Structured filters: categoryId, status, locationId, responsibleUserId, acquisitionYear (YYYY). Archived assets are hidden unless includeArchived=true or status is given explicitly. Accepts a savedViewId (G00 saved-view seam).
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| categoryId | No | ||
| locationId | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| acquisitionYear | No | ||
| includeArchived | No | ||
| responsibleUserId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true, the safety profile is already known from annotations. The description adds meaningful behavioral rules: ordering by number, archived-hidden-by-default, and a savedViewId seam. It goes beyond what annotations convey, though it omits pagination and field-level output details. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences: purpose, filters, archived behavior, saved view. Each sentence adds distinct information with no filler, and the most important information is front-loaded. This is a well-structured, appropriately sized description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 8 parameters and no output schema, so the description carries extra responsibility for explaining behavior and return context. It covers filters and archived defaults but does not mention pagination, page size, or response shape. The cryptic 'G00 saved-view seam' leaves an integration requirement unresolved. Adequate overall, but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it explicitly enumerates the main filter parameters and gives a format for acquisitionYear ('YYYY'). workspaceId is only implied by 'workspace' context, and savedViewId's 'G00 saved-view seam' explanation is cryptic, leaving some parameter semantics under-specified. The description does meaningful work but not complete work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific operation ('List the workspace fixed-asset register') and adds ordering context ('ordered by number'), which clearly identifies the tool's primary purpose. It is distinguishable from asset_get or asset_search by the resource term 'register', though it does not explicitly name sibling alternatives. This is unambiguous but lacks explicit sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on which filters to use and how archived assets behave by default ('Archived assets are hidden unless includeArchived=true or status is given explicitly'), plus the accepted savedViewId. It does not state when to prefer asset_search or asset_get, so exclusions are absent. The guidance is actionable but not comparative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_location_archiveA
Soft-archive a location (active=false): it leaves the transfer/asset pickers but stays resolvable for historical assets. Refused with location_in_use while any non-disposed asset still references it (move those assets away first). Deletion is never offered (a location must stay resolvable for the transfer history that names it). Idempotent: archiving an already-archived location succeeds and writes nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| locationId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden and does so thoroughly. It reveals the soft-archive behavior (active=false), the retention in historical resolution, the exact error condition (location_in_use) and its resolution, the fact deletion is never offered, and idempotency (archiving an already-archived location succeeds and writes nothing). This is exemplary transparency for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each earning its place: core action, side-effect scoping, error precondition, and idempotency guarantee. The most important information is front-loaded, and there is no filler or repetition of schema metadata.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity and the absence of annotations and output schema, the description covers behavior, error conditions, preconditions, and idempotency. It does not state the exact return value or success response format, but for an idempotent soft-archive operation with obvious parameters this is a minor omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. The description alludes to locationId via 'a location' and explains idempotency, which indirectly covers idempotencyKey, but workspaceId is not discussed. Parameter names are self-explanatory and there are only three required fields, so the gap is minor, but the description still does not explicitly map each parameter to its role.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource: 'Soft-archive a location (active=false)', immediately clarifying the tool's core function. It goes further by explaining the effect (leaves pickers, stays resolvable for historical assets) and explicitly distinguishes it from deletion ('Deletion is never offered'), making it clearly distinct from sibling tools like asset_location_update, location_archive, or asset_dispose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives practical context on when the operation is valid and what must happen first: it is refused with location_in_use while any non-disposed asset references the location, and the caller is told to move assets away before archiving. It also clarifies that deletion is never an option. It does not explicitly name sibling alternatives (e.g. asset_location_update or location_archive), but the usage conditions are clear enough for an agent to know when this tool applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_location_createA
Create a workspace fixed-asset location (a physical place assets live: a hall, a floor, a branch). code is 1-30 chars, unique per workspace case-insensitively (duplicate_code); an optional parentId nests it under another location for a light hierarchy (a foreign or missing parent is not_found). Locations are the primary filter dimension of the register and the target of asset_transfer.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | ||
| name | Yes | ||
| parentId | No | ||
| description | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It transparently explains key behaviors and constraints: code uniqueness per workspace case-insensitively with a duplicate_code error, optional parentId with not_found for foreign or missing parents, and the light hierarchy semantics. It does not mention return values or idempotencyKey behavior, but the schema already requires it, and the disclosed error conditions add meaningful value beyond a simple 'create this thing' statement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each dense with relevant information: the core definition, code constraints and error behavior, hierarchy and error handling, and contextual positioning within the asset register. No filler, no repetition, and the most important facts (what it creates) are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create tool with no annotations and no output schema, the description covers the conceptual ground well: what a fixed-asset location is, code uniqueness, optional hierarchy, and how it relates to other asset operations. The main gaps are the lack of explanation for the idempotencyKey (why it is required) and any guidance on the name/description fields. These are not fatal, but the description could be more complete for an agent that needs to construct a valid request.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the schema's lack of parameter documentation. It does explain two key parameters: code (1-30 chars, unique per workspace case-insensitively) and parentId (optional, nests, not_found for invalid parent). However, it provides no guidance on name, description, workspaceId, or the required idempotencyKey. This is partial compensation; the agent still has to infer the purpose of half the parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Create a workspace fixed-asset location') and clarifies what kind of entity it is with concrete examples (hall, floor, branch). It also distinguishes this creation tool from related sibling tools by describing how locations are used ('primary filter dimension of the register and the target of asset_transfer'), so an agent can tell it apart from asset_location_update, asset_location_archive, and asset_location_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong contextual guidance: locations are the primary filter dimension and the target of asset_transfer, which implies when creating a location would be necessary. It does not explicitly name alternatives or state when not to use this tool, but the context is clear enough that an agent can infer the appropriate usage scenario.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_location_getARead-only
Read one fixed-asset location by id, archived or not.
| Name | Required | Description | Default |
|---|---|---|---|
| locationId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes safety. The description adds value by noting 'archived or not', which is a behavioral trait not captured in the annotation—it tells the agent this read will return archived locations as well. This is useful and goes beyond the structured data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that conveys the action, resource, and key nuance (archived or not). Every word earns its place with no waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple get with two parameters, the description is adequate but incomplete. It does not describe the return structure or any error behavior (e.g., what happens if not found). Given no output schema exists, the agent lacks information about the response format. Still, the core purpose and read-only behavior are covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the lack of parameter documentation. It does not: it makes no mention of workspaceId or locationId, their formats, or their roles. The parameter names are self-explanatory, but the description adds zero semantic value, and the coverage gap remains entirely unaddressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'read' and the resource 'fixed-asset location', and specifies it fetches one record by id, including archived ones. This distinguishes it from asset_location_list (multiple) and asset_location_create/update/archive (mutations). The scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when you have a specific location id and need the full record, including archived. It doesn't explicitly name alternatives like asset_location_list, but the 'by id' phrasing strongly signals the intended use case. No exclusions are given, but the context is clear enough for a simple get operation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_location_listARead-only
List the workspace fixed-asset locations, ordered by code. Optional filters: active (true = only live, false = only archived), parentId (a location id, or null for the roots), and a case-insensitive search over code and name. Accepts a savedViewId (G00 saved-view seam).
| Name | Required | Description | Default |
|---|---|---|---|
| active | No | ||
| search | No | ||
| parentId | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true already provided, the description adds meaningful behavioral detail beyond the annotation: it specifies ordering by code, the exact semantics of the active filter (true = live, false = archived), parentId behavior including null for roots, and case-insensitive search over code and name. This helps the agent anticipate tool behavior without overstepping.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler. The verb-resource pair is front-loaded, and each clause adds essential information: ordering, filters, and savedViewId. No redundant phrases or needless elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with no output schema and only a readOnly annotation, the description covers the key input semantics and ordering. It could further mention pagination or result structure, but the absence of an output schema and the straightforward nature of a list operation make this adequate. The savedViewId note is slightly cryptic but acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description carries the full burden of explaining parameters. It does so effectively for all five parameters: workspaceId inferred from 'workspace', active (true/false), parentId (with null semantics), search (case-insensitive over code and name), and savedViewId (G00 saved-view seam). This fully compensates for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific action ('List'), the resource ('workspace fixed-asset locations'), and a distinguishing detail ('ordered by code'). It clearly differentiates from sibling tools like asset_location_get, asset_location_create, and even location_list by specifying 'fixed-asset locations'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the optional filters and the savedViewId seam, giving clear context for when filters might be applied. However, it does not explicitly name alternative tools such as asset_location_get or asset_list, nor does it state when not to use this tool. Usage is implied rather than explicitly guided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_location_updateA
Edit a fixed-asset location through a patch object (name, description, parentId). code is immutable through this verb. Setting a parentId that would close a hierarchy loop is refused with location_cycle; parentId null or "" clears it back to a root location. Only the fields present in patch change.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | No | ||
| locationId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so thoroughly. It discloses immutability of code, the refusal behavior and error code for hierarchy cycles, the effect of null/empty parentId, and that only patch fields are changed. These are exactly the non-obvious behavioral traits an agent needs to know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with zero waste. The core action and key constraints are front-loaded, followed by the most important edge-case and immutability rules. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an update tool with no output schema and no annotations, this description is complete: it tells the agent what can be changed, what cannot, how to clear the parent, what error to expect on a cycle, and that it performs a partial patch. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description compensates well for the patch object by defining the valid fields and explaining the special behavior of parentId, including null clearing to root. It does not describe workspaceId or locationId, but those are unambiguous resource identifiers that need no further explanation. The nuanced patch semantics go far beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Edit') on a specific resource ('fixed-asset location') and the mechanism ('patch object'). It distinguishes itself from sibling tools like asset_location_create and asset_location_archive by clearly framing this as an edit operation, and it further clarifies scope by noting what is immutable ('code').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the usage context clear: this is the tool to modify an existing fixed-asset location's name, description, or parentId. It does not explicitly list alternatives or exclusions, but the edit-vs-create distinction is strongly implied by the verb and sibling tool names, so the context is clear enough for correct selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_maintenance_log_cancelA
Soft-cancel a maintenance log entry entered in error (H08): it sets status=cancelled and stores a required non-empty reason. The row is NEVER hard-deleted (append-oriented, OR 957/958): a cancelled entry stays visible under the list "any" status filter for audit, and is excluded from the default completed-only list and the cost roll-up. Idempotent: cancelling an already-cancelled entry is a no-op that keeps the first reason and writes nothing. Refused with invalid_input for an empty reason and not_found for a foreign id (§H-TENANT). Posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| reason | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so excellently: it discloses soft-delete semantics, that the row is NEVER hard-deleted, visibility under different filters, exclusion from cost roll-up, idempotent no-op behavior, and the exact error conditions (invalid_input, not_found). This is far beyond what the schema or annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but economical: each sentence adds distinct information—action, non-deletion/audit behavior, idempotency, and error contracts. The core purpose is front-loaded, and no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter mutation tool with no annotations and no output schema, the description is remarkably complete. It covers what happens to the row, how it appears in lists, idempotency, failure modes, and confirms no ledger posting. An agent can invoke this tool correctly without further documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains that reason must be non-empty, that id refers to a foreign maintenance log entry, and that idempotency is keyed behavior. It does not explicitly describe workspaceId, but the tenant scoping is implied via §H-TENANT and the foreign-id error. Substantial value added over the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Soft-cancel'), resource ('maintenance log entry'), and condition ('entered in error'), and distinguishes the operation from hard deletion. The H08 reference and status change make the tool's purpose unmistakable even among the asset_maintenance_log_* siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use: entries entered in error should be soft-cancelled, and the description explains why (append-oriented, audit visibility). It does not explicitly name an alternative or say 'use update instead of cancel,' but it is clear enough for an agent to identify the correct operation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_maintenance_log_createA
Record a completed maintenance event against a fixed asset (H08): what was done, when, by whom and at what cost. NON-POSTING: it writes ONE asset_maintenance_log row (status=completed) and posts NO GL journal, so the captured cost is descriptive TCO metadata (feeding H09), never an expense or capitalisation booking (use the J expense or OP16 capitalisation flow for that). assetId must name an asset in this workspace that is NOT archived (asset_archived; history stays readable but new entries are blocked). maintenanceType is one of corrective | preventive | inspection | calibration | upgrade | other (invalid_maintenance_type). logDate is a required ISO date no more than 30 days in the future (missing_log_date / log_date_too_far). title is 1-200 chars. Every cost field is optional, integer Rappen, >= 0 (invalid_cost, never a float); when partsCostRappen and labourCostRappen are both given, a supplied costRappen total must equal their sum, otherwise the total is derived from them. Refused with not_found for a foreign asset (§H-TENANT). Idempotent on idempotencyKey: a replay returns the original log and writes no second row.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| title | Yes | ||
| assetId | Yes | ||
| logDate | Yes | ||
| costRappen | No | ||
| description | No | ||
| workspaceId | Yes | ||
| externalParty | No | ||
| idempotencyKey | Yes | ||
| maintenanceType | Yes | ||
| partsCostRappen | No | ||
| labourCostRappen | No | ||
| linkedDocumentId | No | ||
| externalReference | No | ||
| performedByUserId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses non-posting behavior, idempotency on idempotencyKey (replay returns original log, no second row), and validation rules (cost derivation, not_found for foreign asset). This is strong coverage for a mutation tool with zero annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded with the core purpose. Every sentence adds value, though it is long and packed with many details, which slightly reduces readability. Still, it avoids fluff and is structured logically from general to specific.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 15 parameters, 0% schema coverage, and no output schema or annotations, this description is comprehensive. It covers all validation rules, error codes, idempotency, and cost aggregation logic. An agent has everything necessary to call it correctly, including distinguishing from sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are 15 parameters, so the description is the only source. It adds semantics for many parameters: assetId (working, not archived), maintenanceType (enum values and error code), logDate (required ISO, 30-day future limit), title (1-200 chars), cost fields (optional, integer Rappen, >=0, derivation rules), idempotencyKey (idempotent behavior). It covers all the critical ones with error codes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool records a completed maintenance event (verb 'record' + specific resource 'maintenance log'), and explicitly notes it is NON-POSTING, distinguishing it from financial posting tools. It covers all key aspects (what, when, by whom, cost) and differentiates from siblings like asset_depreciation_run_post.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use (recording a completed maintenance event) and when not (for expense/capitalisation booking, directing to J expense or OP16 capitalisation flow). It also gives clear conditions like asset must be in workspace and not archived, and logDate constraints.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_maintenance_log_getARead-only
Read one maintenance log entry by id (H08), completed or cancelled. A foreign id is not_found (§H-TENANT), never a leak of another workspace row.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description doesn't need to repeat that. It adds the error behavior: 'A foreign id is not_found (§H-TENANT), never a leak of another workspace row.' This is valuable context about error handling and tenant isolation, beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no redundancy. The main action is first, and the error note is brief. Highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core purpose and error behavior. However, it doesn't clarify the meaning of workspaceId or the exact statuses that can be read (is it restricted to completed/cancelled or can it read all?). There is no output schema, so the return format is not described. For a simple get, it's mostly complete but has minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 0% description coverage, so the description must explain the parameters. It mentions 'by id' which refers to the id parameter, but does not explain workspaceId at all. The description fails to compensate for the lack of schema descriptions, leaving workspaceId ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb 'Read' and resource 'maintenance log entry' with scope 'by id'. Distinguishes from sibling tools like asset_maintenance_log_list (which lists multiple) and asset_maintenance_log_create (which creates). Also specifies statuses 'completed or cancelled'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states the operation (read by id) and the required parameters (workspaceId and id) are evident from the schema, but it doesn't explicitly mention when to use this vs list or other alternatives. However, the name and description make it obvious that this is for fetching a single entry, so the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_maintenance_log_listARead-only
List maintenance log entries (H08), newest first (by logDate then created_at). Optional filters: assetId (scope to one asset, the usual timeline read), maintenanceType, status (completed [default] | cancelled | any), fromDate / toDate (an ISO log-date window), hasCost (true = only entries with a captured cost, false = only those without), performedByUserId, and a case-insensitive search over title, description, externalParty and externalReference. Returns items, total, and totalCostRappen: the SUM of cost_rappen over the COMPLETED entries in the result (a cancelled or costless entry contributes 0), which is the asset TCO roll-up. Accepts a savedViewId (G00 saved-view seam). Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| search | No | ||
| status | No | ||
| toDate | No | ||
| assetId | No | ||
| hasCost | No | ||
| fromDate | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| maintenanceType | No | ||
| performedByUserId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses sort order, default status behavior, the exact semantics of totalCostRappen (sum over completed entries only), and the saved-view seam. This significantly enriches the agent's understanding of the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but densely informative and well-structured: purpose and ordering first, then filters, then return semantics. Every sentence adds operational value, and the length is justified by the large parameter surface.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description correctly explains the return shape (items, total, totalCostRappen) and the special aggregation rule. Combined with parameter semantics and behavioral transparency, the agent has enough context to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description carries the full burden and does so thoroughly: each major filter is explained, including status values, date-window semantics, hasCost behavior, and the fields covered by the case-insensitive search. Only workspaceId is left implicit, which is reasonable for a required workspace-scoped parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb and resource ('List maintenance log entries (H08)') with explicit ordering and scope. It is clearly distinguished from related sibling tools like asset_maintenance_log_get or asset_maintenance_log_create by focusing on listing with filters.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool, including the intended asset timeline read and a default status of 'completed'. It does not explicitly name alternatives or state when not to use it, but the filter semantics and 'usual timeline read' hint provide strong practical guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_maintenance_log_updateA
Correct a recent maintenance log entry (H08) through a patch object: title, description, costRappen, partsCostRappen, labourCostRappen, externalReference, notes. Identity fields (assetId, logDate, maintenanceType, the performer) are immutable after create; only the descriptive/cost fields present in patch change. Refused with log_locked once the entry is older than the workspace soft-edit window (default 90 days from creation: after it only cancel is allowed, protecting the historical record), log_cancelled if the entry is already cancelled (correct it with a fresh entry, not an edit), invalid_cost for a negative or non-integer cost, and not_found for a foreign id (§H-TENANT). Posts nothing. Idempotent on idempotencyKey.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| patch | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden – and it delivers. It discloses immutable identity fields, the 90-day soft-edit window, idempotency on idempotencyKey, 'Posts nothing,' and the exact error cases. This goes well beyond what a generic update would imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence contributes: purpose, editable scope, immutability, error conditions, cancellation alternative, idempotency. There is no fluff or restatement of the schema, and the core action is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 parameters, a nested patch object, no annotations, and no output schema, the description fully covers invocation rules, failure modes, and side effects. It gives an agent all needed to call it correctly, even without seeing sibling tool definitions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, but the description compensates by listing which patch fields are editable, which are immutable, and the validation rules for costs (negative/non-integer → invalid_cost). It also explains the role of idempotencyKey, giving agents semantic understanding that the raw schema alone lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Correct') and resource ('maintenance log entry (H08)') with a patch object, and enumerates the exact editable fields. It distinguishes itself from the cancellation and creation alternatives by noting that cancelled entries need a fresh entry and locked entries only allow cancel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly provides when-to-use vs. alternatives: 'only cancel is allowed' after the soft-edit window, 'correct it with a fresh entry, not an edit' for cancelled entries. Also lists error conditions (log_locked, log_cancelled, invalid_cost, not_found) so an agent can decide whether to retry or switch tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_nbv_summaryARead-only
A dimensional NBV / cost roll-up (H09, US-H09.7): grouped rows (key, count, sumCostRappen, sumAccumRappen, sumNbvRappen) plus a grandTotal. groupBy is category | location | status | method (cost_center is not an asset-master dimension in Phase 1: invalid_input). Non-disposed assets only by default; includeDisposed:true folds disposed assets in. Optional register filter (the asset_register_report filter shape) narrows the base set. Pure read over the asset master, §H-TENANT.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No | ||
| groupBy | Yes | ||
| workspaceId | Yes | ||
| includeDisposed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, and the description reinforces this with 'Pure read over the asset master'. Beyond that, it discloses meaningful behavioral details: the default non-disposed filter, the effect of includeDisposed, and the rejection of cost_center as an invalid groupBy value. This enriches the agent's understanding of the tool's behavior beyond what annotations state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is five dense sentences with zero fluff. It front-loads the purpose ('dimensional NBV / cost roll-up'), then packs the output row fields, grouping constraints, default filtering, filter reference, and read-only nature into a compact structure. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description lists all row fields and the grandTotal, giving the agent a clear expectation of the response shape. It covers the default behavior, valid/invalid groupBy values, and the filter shape reference. It does omit detailed filter property semantics必胜, but defers to a sibling tool's filter shape, which is acceptable. The §H-TENANT and report code references may assume system familiarity but are not required for basic invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates significantly. It defines the allowed groupBy values and marks cost_center as invalid, explains includeDisposed's default behavior, and references the asset_register_report filter shape to describe the filter parameter. WorkspaceId is left generic, which is acceptable since it is a standard parameter. While the filter properties are not individually described, the reference to a sibling's filter shape is a useful shortcut.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies this as a dimensional NBV/cost roll-up, lists the exact output row fields and grandTotal, and specifies the supported groupBy dimensions (category, location, status, method). It also explicitly calls out that cost_center is not a valid dimension, distinguishing this summary tool from other asset tools like asset_list or asset_search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: when a grouped NBV/cost roll-up is needed, with specific grouping options and a default excluding disposed assets. It also explains how to include disposed assets and that an optional filter can narrow the base set. However, it does not explicitly name alternative tools or state when not to use them, relying on the summary nature to imply selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_opening_balanceA
Seed the opening cost and accumulated depreciation of a fixed asset (H07, US-H07.5): for an asset that already exists in the real world at migration, it records the historical figures so the identity invariants and the recon report hold from the opening period. In ONE atomic transaction it posts ONE balanced GL journal via A02 (Dr the asset cost account for costRappen, Cr the accumulated-depreciation account for accumulatedDeprRappen when > 0, and Cr the offsetAccountId equity/opening account for the net book value cost minus accumulated) and writes ONE append-only asset_transaction of type=opening naming that journal, then moves the asset from draft to active with that baseline (which trips the H01 financial-field lock). offsetAccountId is required only when costRappen > accumulatedDeprRappen. Refused with not_found (foreign asset, §H-TENANT), already_acquired / already_opened (the asset already carries a financial event), invalid_input (bad date, non-positive cost, accumulated outside 0..cost), invalid_offset_account, and period_locked. Idempotent on idempotencyKey: a replay returns the original objects (asset, transaction, journalEntry) and posts no second journal and appends no second row.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | ||
| assetId | Yes | ||
| costRappen | Yes | ||
| description | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| offsetAccountId | No | ||
| accumulatedDeprRappen | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so thoroughly. It discloses atomicity, exact journal postings via A02, the append-only transaction row, the draft-to-active state transition, the H01 financial-field lock, the full error contract, and idempotency semantics. This goes well beyond the structural fields and gives an agent a realistic model of side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Although long, the description is dense and front-loaded: purpose first, then accounting mechanics, state transition, error cases, and idempotency. Every clause adds operational information needed to invoke the tool safely. The structure mirrors the invocation-order concerns an agent would need, so the length is justified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex financial tool with 8 parameters, no annotations, and no output schema, the description is remarkably complete. It covers what the tool does, how the journal is balanced, when the offset account is required, what errors can occur, and what idempotent replay returns. No critical behavioral dimension needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains costRappen, accumulatedDeprRappen, offsetAccountId, and idempotencyKey, including the condition under which offsetAccountId is required. It also encodes validation rules ('non-positive cost, accumulated outside 0..cost'). It does not separately define workspaceId, assetId, and description, but those are largely inferable from the resource context and naming.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Seed the opening cost and accumulated depreciation of a fixed asset'. It clearly scopes the tool to migration-time assets that already exist in the real world, and explains the downstream invariants (identity invariants and recon report). This distinguishes it from asset acquisition and other asset lifecycle tools without needing to inspect siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes the intended scenario: 'for an asset that already exists in the real world at migration'. It also lists refusal conditions including already_acquired/already_opened, which implicitly tells agents when not to call it. However, it does not explicitly name alternative tools such as asset_acquire for new assets, so the routing guidance is contextual rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_reconciliation_checkARead-only
The HARD fixed-asset reconciliation check for a period (H07, US-H07.3), the one period-close and agents invoke. It reconciles the whole workspace at the period end (YYYY-MM) and returns { status: "balanced", accounts } when every control account nets to 0 difference, or the structured reconciliation_drift error naming the offending accounts and their sub-ledger / GL / delta amounts when any does not. Period close may hard-lock only on the balanced answer; a drift blocks it until the difference is explained or corrected through the ordinary reversing + correcting flow (the reconciliation itself posts nothing). Read-only, §H-TENANT.
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description reinforces this by stating it posts nothing. It adds extra behavioral context: it can hard-lock period close on balance, returns a structured error on drift, and the drift must be resolved before close. This goes beyond the annotation by explaining the operational consequence and the correction path.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but informative, front-loading the critical 'HARD fixed-asset reconciliation check' and the fact it's the one invoked. It is a bit long but each clause adds necessary detail about behavior and consequences. It could be slightly trimmed but is well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only check with no output schema, the description covers the success and error return shapes, the period format, and the impact on period close. It lacks explicit mention of required parameters but the context makes them inferable. Given the complexity of the tool and minimal schema, it is quite complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explicitly explain the parameters. However, it mentions 'period (YYYY-MM)' and 'whole workspace', implying workspaceId and period. The period format is given, which adds value, but workspaceId semantics are left vague. With 0% coverage, a 3 is appropriate as the description partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is the hard fixed-asset reconciliation check for a period, identifies the specific codes (H07, US-H07.3), and explains it reconciles the whole workspace at period end. It distinguishes itself from other asset tools by being 'the one period-close and agents invoke' and by describing the exact output on success and failure.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use (period close) and what happens on a balanced vs. drift result, and mentions that a drift blocks period close until corrected through the ordinary reversing + correcting flow. It also notes the reconciliation itself posts nothing, which is a clear indication of when not to use it for posting actions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_reconciliation_reportARead-only
The OP11 reconciliation report for fixed assets (H07, US-H07.2): one row per control account (cost and accumulated depreciation), each with the sub-ledger total (summed from asset_transaction), the posted GL balance of the same account, the delta (0 when balanced), a balanced | drift status, and the contributing-asset drill-down. Cut-off: as_of (an ISO date) takes precedence over period (YYYY-MM, whose month-end is the cut-off); with neither it is as-of today. Optional accountIds filter. A disposed asset drops out of the open totals because its disposal cleared both sides; a journal posted directly against a control account with no asset_transaction shows as drift (the detection mechanism, never auto-corrected). Read-only, §H-TENANT.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| period | No | ||
| accountIds | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses material behaviors: disposed assets drop out, direct journal postings appear as drift, and the tool never auto-corrects. It also adds the tenant scope (§H-TENANT) and the precedence logic, giving the agent a robust behavioral model.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense yet efficiently structured: report definition first, then parameters, then edge casesholistically. Every sentence adds value, and the read-only note is front-loaded in the final sentence but already implied by the annotation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a report tool with no output schema, the description covers enough to call it correctly: required workspaceId (via schema), optional filters, cut-off semantics, row structure, and detection behavior. The edge cases around disposal and direct journal postings prevent false expectations, making it complete for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does so for asOf (ISO date), period (YYYY-MM, month-end), and accountIds (optional filter), but it omits workspaceId, which is required yet not explained beyond the schema. Still, the most ambiguous parameters are well documented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific report (OP11 reconciliation report for fixed assets) and clearly states its output: one row per control account with sub-ledger total, GL balance, delta, status, and drill-down. It is immediately distinguishable from sibling tools like asset_reconciliation_check or asset_ledger_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when this read-only report is useful (drift detection) and how its cut-off is determined, with clear precedence rules for as_of vs period. It does not explicitly name alternatives or say when not to use it, so it falls short of a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_register_reportARead-only
The complete, filterable fixed-asset register (H09, US-H09.1): one row per asset (number, name, category code/name, status, acquisitionDate, acquisitionCostRappen, accumulatedDeprRappen, netBookValueRappen, residualValueRappen, method, usefulLifeMonths, location, responsible, serial/barcode, lastDepreciationPeriod, disposedAt) with a totals footer (count, sum cost, sum accum, sum NBV) computed over the WHOLE filtered set, not just the returned page. filter carries status[] (default hides archived), categoryIds[], locationIds[], acquisitionYearFrom/To (YYYY), costMinRappen/costMaxRappen, responsibleUserId and a case-insensitive q over number/name/serial/barcode/notes. Optional columns subset (id/number always kept), sort ({ field, dir }) over a whitelisted column, and cursor/limit pagination (default 500, max 5000; follow nextCursor until absent). An empty match is a success ({ rows: [], zero totals, message: "no_assets_matching" }), never an error. Pure read, §H-TENANT.
| Name | Required | Description | Default |
|---|---|---|---|
| sort | No | ||
| limit | No | ||
| cursor | No | ||
| filter | No | ||
| columns | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses many non-obvious behaviors: totals computed over the whole filtered set, not just the returned page; empty matches return success with a message; default status hides archived; pagination via cursor until absent; default 500/max 5000 limit; case-insensitive q; and pure read with §H-TENANT. This is rich behavioral context beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause earns its place, moving from purpose, to output fields, to totals behavior, filters, pagination, edge case, and read-only guarantee. It is front-loaded with the report identity and scope before diving into details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers output fields, filter semantics, pagination, empty-match behavior, and the read-only nature, which is sufficient for a complex report with no output schema. It does not enumerate status values or list whitelisted sort fields, which are minor gaps given the schema has no enums, but the overall context is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description fully compensates: it explains filter semantics (status defaults, categoryIds, locationIds, acquisition year range, cost range, responsibleUserId, case-insensitive q over specific fields), columns subset behavior (id/number always kept), sort over a whitelist, and cursor/limit pagination. The only minor gap is the exact sort field whitelist values, but the field list is already enumerated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'complete, filterable fixed-asset register (H09, US-H09.1)', identifying the exact report code and stating 'one row per asset' with a full enumerated field list. It is easily distinguishable from sibling asset tools like asset_list or asset_search by its report scope and totals footer.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly establishes when this tool is appropriate: when a filterable register with totals over the whole filtered set is needed. It does not explicitly name sibling alternatives or state when not to use it, but the context is unambiguous enough for an agent to select it correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_searchARead-only
Full-text search the register over number, name, serial number, barcode and notes (case-insensitive substring). An empty query returns no rows.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true, which covers safety. The description adds valuable behavioral details beyond that: it specifies search fields, case-insensitive substring matching, and the edge case of empty query returning no rows. This enriches the agent's understanding without contradicting annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence that front-loads the core action ('Full-text search') and packs in the key details without any wasted words. It is perfectly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and minimal annotations, the description covers the search behavior and an important edge case, but it omits expected return format, pagination/limit behavior, and ordering. It also does not clarify the role of workspaceId. For a search tool, these gaps are moderate; the description is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for both parameters. It explains the meaning of 'query' implicitly (the search text over the listed fields), but it does not mention 'workspaceId' at all, leaving its purpose (likely scoping the search) undocumented. Only one of two parameters is partially explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it performs full-text search over specific fields (number, name, serial number, barcode, notes) with case-insensitive substring matching. This distinguishes it from other asset tools like asset_list (likely listing) and asset_get (fetch by ID). The verb 'search' and resource 'register' are explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (when you need to find assets by text across multiple fields), but it does not explicitly contrast with alternatives like serial_search, lot_search, or asset_list. No 'when not to use' guidance is given, but the context is clear enough for an agent to infer.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_transaction_getCRead-only
Read one fixed-asset sub-ledger transaction by id (H02), including the GL journal entry it posted and its source-document link.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers the safety profile, so the description does not need to restate that. It adds value by disclosing that the response includes the GL journal entry and source-document link, which is useful behavioral context about return content. However, it does not describe any other behavior (e.g., what happens if the ID is invalid, pagination, or rate limits), and the annotation is not contradicted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly worded sentence that front-loads the action and resource, then lists the included return data. It contains no filler or redundancy, making it highly efficient and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple getter with two parameters and no output schema, the description conveys the core purpose and return contents. However, it omits any explanation of the parameters and the H02 reference is left unexplained, which could confuse an agent unfamiliar with the domain. It is adequate but not comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage, and the tool description does not explain what 'id' or 'workspaceId' mean or how they should be formatted. There is no guidance on the ID's structure (e.g., H02 identifier) or the workspace scope. The description fails to compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (Read), the resource (fixed-asset sub-ledger transaction), and the scope (by id). It also mentions what is included in the result (GL journal entry, source-document link), which gives a precise picture. However, it does not explicitly differentiate from sibling tools like asset_transaction_list, though the 'by id' phrasing implies a single record lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives. It does not mention that this is for retrieving a single transaction when you have its ID, nor does it contrast with list or search tools. An agent would have to infer usage from the 'by id' wording, which is not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_transaction_listARead-only
List the financial-event history of one asset (H02), ordered oldest first: every asset_transaction row with its type, date, capitalised amount (deltaCostRappen), the GL journal entry it posted (journalEntryId) and any source-document link. Optional type filter (acquisition | additional_capitalisation).
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | ||
| assetId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true; the description adds valuable behavioral context: result ordering, asset scoping, the optional type filter values, and the exact fields returned. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single focused sentence that front-loads purpose and scope, then packs the return fields and filter option efficiently. There is no redundant or filler wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with no output schema, the description provides enough context: what is listed, for which entity, in what order, and what optional filtering exists. The absence of pagination or limit details is a minor gap but not enough to lower the score further.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description partially compensates by clarifying assetId as the target asset and type as an optional filter with allowed values. However, workspaceId is not mentioned and there are no parameter-level explanations, leaving some burden unmet.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('List'), a clear resource ('financial-event history of one asset'), and details exactly what is returned (type, date, capitalised amount, journal entry ID, source-document link) and the ordering (oldest first). It is clearly distinguishable from siblings like asset_transaction_get or asset_ledger_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes it clear that this tool is for listing financial-event history for one asset, and the optional type filter gives a selection criterion. However, it does not explicitly mention alternatives or state when not to use this tool versus sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_transferA
Transfer one or many active fixed assets to a new location and/or a new responsible person (custodian). NON-POSTING by design: it writes one immutable asset_transfer history row per asset (capturing the old and new location/responsible, the effective date and an optional reason) and updates each asset location_id / responsible_user_id, but changes NO financial field and creates NO journal entry. assetIds is 1..N (bulk supported, all-or-nothing); at least one of toLocationId / toResponsibleUserId must be supplied (nothing_to_transfer otherwise). Refused with location_inactive (target location archived), asset_not_transferable (any selected asset is disposed or archived, listing the offenders), or not_found (a foreign asset or location, §H-TENANT). Idempotent on idempotencyKey: a replay returns the original {transactions, assets, summary} and writes no second history row.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| assetIds | Yes | ||
| workspaceId | Yes | ||
| toLocationId | No | ||
| effectiveDate | Yes | ||
| idempotencyKey | Yes | ||
| toResponsibleUserId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so thoroughly. It discloses the non-posting design, the immutable history row write per asset, the exact fields updated (location_id / responsible_user_id), the all-or-nothing bulk behavior, the idempotency behavior on idempotencyKey, and the specific error conditions. This is far beyond what a typical description provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-rich, with the core purpose front-loaded in the first sentence. It packs a lot of behavioral detail into a compact paragraph. It loses one point because the density of error codes and edge cases makes it slightly harder to parse quickly, but every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description is remarkably complete. It covers the operation's scope, side effects, constraints, error conditions, idempotency, and return behavior (replay returns original {transactions, assets, summary}). An agent has everything it needs to decide whether to call this tool and what to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the schema's lack of parameter documentation. It does explain the key parameters: assetIds (1..N, bulk supported), toLocationId / toResponsibleUserId (at least one required), and idempotencyKey (idempotent replay behavior). However, it does not explain workspaceId, effectiveDate, or reason in detail, though their purpose is fairly inferable from the schema and context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Transfer'), a precise resource ('active fixed assets'), and the exact scope of the operation ('to a new location and/or a new responsible person (custodian)'). It also distinguishes itself from related asset tools by explicitly noting it is NON-POSTING and writes no journal entry, which separates it from asset_dispose, asset_acquire, and other asset mutation tools in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use guidance: it is for transferring assets to a new location and/or custodian, and it explicitly states what it does NOT do (no financial fields, no journal entry), which helps an agent avoid using it for financial postings. It also names specific error conditions (location_inactive, asset_not_transferable, not_found) that signal when the tool is being used inappropriately, and mentions the sibling asset_transfer_history for related history reads.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_transfer_historyARead-only
List the complete transfer history of one asset (H05), ordered oldest first: every asset_transfer row with its effective date, from/to location, from/to responsible person and reason. Read-only; no financial figure is ever part of a transfer.
| Name | Required | Description | Default |
|---|---|---|---|
| assetId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces it while adding 'no financial figure is ever part of a transfer' – an extra behavioral guarantee beyond the annotation. It also discloses ordering and included fields. No contradiction exists; the description enriches the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence with the primary action front-loaded. Every phrase adds value: the scope, ordering, field list, and read-only note. No wasted words, and it is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with two parameters and no output schema, the description covers the returned fields and read-only nature. However, it omits any parameter semantics, potential pagination, or response size considerations. It is functional but leaves gaps that could matter for large histories.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, and the description does not explain either workspaceId or assetId. While it mentions 'one asset', it does not map that to the assetId parameter or clarify the role of workspaceId. The description fails to compensate for the lack of schema documentation, leaving parameter semantics completely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists the complete transfer history of one asset, specifies the ordering (oldest first), and enumerates the fields returned (effective date, from/to location, from/to responsible person, reason). It differentiates from related tools like asset_transfer (which likely performs transfers) and asset_transaction_list by focusing specifically on transfer history. The read-only note and exclusion of financial figures further sharpen the purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for retrieving transfer history but does not explicitly contrast with alternative tools or state when to choose this over others. There are no exclusions or named siblings, though the read-only and non-financial clarifications hint at scenarios where financial tools would be inappropriate. This is adequate but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
asset_updateA
Edit an asset through a patch object. Descriptive fields (name, description, location, responsible, serial, barcode, manufacturer, model, warranty, notes) change freely while the asset is non-terminal. The financial baseline (acquisitionDate, acquisitionCostRappen, the depreciation trio, the three GL accounts) may change ONLY while the asset is still draft: once a posted acquisition or depreciation exists, any such field in the patch is refused with financial_fields_locked and nothing is written (use a correction flow instead). A disposed or archived asset is refused with asset_terminal.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | No | ||
| assetId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the full burden of behavioral disclosure. It does so by explaining that financial fields are refused with 'financial_fields_locked' and nothing is written, and that disposed/archived assets are refused with 'asset_terminal'. It also states that descriptive fields change freely, implying non-destructive behavior for those fields. However, it does not describe the success response, idempotency behavior, or whether the operation is atomic, leaving some gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two long sentences but is well-structured: the first sentence introduces the tool and the freely editable fields; the second explains the restrictions and error codes. It is front-loaded with the primary purpose and includes only necessary detail. While dense, it avoids redundancy and maintains focus, making it reasonably concise for the complexity involved.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity—a patch with 19 fields, state-dependent constraints, and no output schema or annotations—the description covers the core behavior but leaves gaps. It does not explain the meaning of each parameter beyond broad categories, does not mention what the tool returns on success, and does not detail the idempotencyKey's role. It is sufficient for a basic understanding but incomplete for an agent to fully predict all outcomes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It does group the patch fields into descriptive vs. financial categories and explains the constraints on each, which adds meaning beyond the schema's bare types. However, it does not explain the purpose or format of parameters like workspaceId, assetId, idempotencyKey, or the exact meaning of 'depreciation trio' and 'three GL accounts' beyond naming them. It adds some value but is not comprehensive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Edit an asset through a patch object.' It specifies the verb (edit), the resource (asset), and the mechanism (patch), making it distinct from sibling tools like asset_create, asset_get, or asset_archive. It also enumerates the editable fields, which removes any ambiguity about what this tool is for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong usage context by distinguishing when fields are freely editable (descriptive fields for non-terminal assets) and when they are locked (financial fields unless the asset is draft). It also directs the user to 'use a correction flow instead' for locked financial fields, effectively naming an alternative. However, it does not explicitly list sibling tools like asset_create or asset_archive as alternatives, so it is not fully explicit about when to choose this tool over others.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
attach_receiptA
Record which Beleg belongs to a bill (OR 958f: the receipt is kept for ten years, and this is what says WHICH one). A reference, not a file: file storage is E00 and A17 stores the pointer. Permitted on a posted bill, because a Buchungsbeleg is filed after the booking at least as often as before it; it is the only column a posted bill will let you change.
| Name | Required | Description | Default |
|---|---|---|---|
| receiptRef | Yes | ||
| workspaceId | Yes | ||
| vendorBillId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses meaningful behavior: the tool mutates a posted bill (unusual), records a pointer rather than file content, and cites the ten-year retention rule (OR 958f). It does not disclose overwrite semantics, validation failures, return values, or the role of the idempotency key.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences with the core purpose front-loaded. Every clause earns its place (purpose, reference-not-file distinction, posted-bill allowance). The density of accounting jargon (OR 958f, E00, A17, Buchungsbeleg) slightly hurts parseability, but there is no waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Four required parameters with no annotations and no output schema puts a heavy burden on the description. It covers purpose, legal rationale, storage separation, and posted-bill behavior well, but omits return/confirmation behavior, failure conditions, and the meaning of two parameters. Adequate for a simple reference-linking tool, with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains receiptRef (a pointer/reference, not a file) and implicitly vendorBillId (the bill being attached to), but says nothing about workspaceId or idempotencyKey. Two of four parameters remain semantically unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action and resource: 'Record which Beleg belongs to a bill'. It also distinguishes itself from file storage ('A reference, not a file'), which differentiates it from sibling file-related tools. However, it relies on domain jargon (Beleg, Buchungsbeleg) and doesn't name a sibling tool explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear when-to-use context: 'Permitted on a posted bill' and explains the reference-vs-file boundary ('A reference, not a file: file storage is E00 and A17 stores the pointer'). It gives the exception that this is 'the only column a posted bill will let you change'. It lacks explicit when-not-to-use phrasing and named alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
attention_listARead-only
Page one queue or one area of Pendenzen: the same providers, the same capability filter and the same collapse and ranking as attention_summary, so the two faces never describe a queue differently. Returns { computedAt, items:[AttentionItem], nextCursor?, failed:[queueId] }. Filter by queueId (a single queue), area ('bank' | 'sales' | ...), or urgency ('overdue' | 'due' | 'open'); page with limit (default 20) and the opaque cursor from a prior nextCursor. A denied queue is absent, a thrown provider is named in failed[], and nothing here clears, dismisses or approves anything: every AttentionItem carries its decisionOptions[] (the owning surface's exits as existing write verbs with their fixed input and per-item idempotencyKey, see attention_summary), its suggestedInvoiceId, reasonCode/reasonKey and consequenceKey/consequence, and a deepLink descriptor into its owning surface for the hard case. Reads only.
| Name | Required | Description | Default |
|---|---|---|---|
| area | No | ||
| limit | No | ||
| cursor | No | ||
| queueId | No | ||
| urgency | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, it explicitly says 'Reads only' and 'nothing here clears, dismisses or approves anything'. It also discloses failure behavior: denied queues are absent, thrown providers appear in failed[]. This adds meaningful behavioral context beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but information-dense; every clause adds parameter semantics, return shape, failure semantics, or item contents. It is front-loaded with the main purpose and structured clearly enough for a complex read tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description correctly explains the return shape: computedAt, items, nextCursor?, failed[]. It also details filter options, paging, failure behavior, and key fields on AttentionItem, referencing attention_summary for decisionOptions. This is sufficient for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the burden, and it covers most parameters well: queueId, area, urgency, limit (default 20), and cursor (opaque, from prior nextCursor). It does not explain the required workspaceId, which is a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('page') and resource ('one queue or one area of Pendenzen'), and explicitly ties itself to attention_summary as the two faces of the same queue. This clearly differentiates it from a generic list and from sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description names attention_summary as the related surface and explains that it shares providers, capability filtering, collapse, and ranking. It gives clear filtering and paging context, but it does not explicitly state when to choose attention_list over attention_summary or include exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
attention_summaryARead-only
Pendenzen (the attention hub): everything waiting for a decision from the caller, in ONE call, composed live from the module queues the product already owns and never cached (pure read, G15 posts nothing and owns no tables). Returns { computedAt, visibleQueues, total, incomplete, queues:[{queueId, area, count, topUrgency}], top:[AttentionItem], failed:[queueId] }. 'top' is the ranked, same-entity-collapsed leading rows the human screen shows (urgency overdue>due>open, then a fixed queue rank, then since), so an agent and the screen read one payload. Each queue's 'count' is a true COUNT, never a truncated list length. HONESTY CONTRACT an agent may quote: a queue whose read capability the caller lacks is ABSENT from queues[] and top[] entirely (not a zero); a provider that throws is NAMED in failed[] with incomplete:true and contributes no count and no row; and 'total' is null (never 0) when visibleQueues is 0, because an actor told nothing was not told a total. topLimit defaults to 5. Registered queues: agent_action (A35 pending Vorschläge -> /agent), qr_match (A21 unmatched incoming credits -> /reconciliation), review_flag (A25 flagged postings -> /journal) and dunning_run (A15 proposed Mahnläufe -> /dunning). THE ROW CARRIES ITS DECISION (F-01): each AttentionItem names the exits its owning surface offers as decisionOptions[] ({id, verb, labelKey, role, input, humanConfirm?, reasonField?, capability?, deepLink?}), where verb is an EXISTING write (apply_qr_match, override_qr_match, approve_drafted_action, reject_drafted_action, issue_dunning_run, approve_entry) and input is the whole fixed input beyond workspaceId incl. a deterministic per-item idempotencyKey (replay-safe: call it twice, it writes once); plus suggestedInvoiceId (A21's live-scored invoice, null when none), reasonCode/reasonKey (why it is pending), consequenceKey/consequence (the D118 C4 sentence of the write from the dial map; null when the verb is not dial-governed) and proposedBy (the raw actor id of the proposer or reviewer, for the self-approve check) with proposedByKind ('agent' | 'studio' | 'member' | 'unknown') and proposedByName (a member's display name, else null), so a surface names the actor in words and never prints an id. The hub mints no verb and clears nothing itself: an agent acts by calling the option's verb with its input, under that verb's own gate and dial.
| Name | Required | Description | Default |
|---|---|---|---|
| topLimit | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the readOnlyHint annotation. It details the honesty contract (absence vs zero, failed providers, total null semantics), the exact return structure, the meaning of each field, the default topLimit, the registered queues, and the decision options with idempotency guarantees. It also clarifies that the tool mints no verbs and clears nothing. This is exemplary behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely verbose, running several hundred words, and is not front-loaded. The essential summary is in the first sentence, but the rest is a dense technical dump covering return shapes, honesty contracts, registered queues, and decision option semantics. While all information is relevant, it is not concise; a more structured layout with bullet points would improve scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and the absence of an output schema, the description is remarkably complete. It covers the return payload, the honesty rules, the meaning of every field, the queue registry, the decision options with verbs and input construction, and the idempotency behavior. An agent has everything needed to call this tool correctly and interpret its results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description must explain the parameters. It mentions that 'topLimit defaults to 5' and implies workspaceId is the required workspace context, but it does not describe what topLimit limits (presumably the size of top[]) or any constraints. It adds some value beyond the bare schema but leaves the parameters under-explained given the tool's complexity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a crisp statement of purpose: 'Pendenzen (the attention hub): everything waiting for a decision from the caller, in ONE call, composed live from the module queues the product already owns and never cached.' It names the resource, the scope, and the aggregation behavior, and explicitly notes it is a pure read. This clearly distinguishes it from any sibling that might offer a simpler list or detail view.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus alternatives, nor does it mention any exclusions or when not to use it. It implies it is the single-call hub for pending decisions, but the sibling list includes attention_list and many action verbs; there is no guidance on when an agent should choose attention_summary over those. This leaves the selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
balance_sheetARead-only
Bilanz (balance sheet) as of a date, grouped into the FIRST LEVEL of OR Art. 959a: Umlaufvermögen and Anlagevermögen on the Aktiven, kurzfristiges Fremdkapital, langfristiges Fremdkapital and Eigenkapital on the Passiven, each section carrying its statutory heading in de/fr/it/en, plus one residual section per side (the "weitere Positionen" of OR Art. 959a Abs. 3) so no account can be dropped. Under each section the accounts are listed individually, ordered by account number. Every line is positive on its own side. The running result IS carried as an Eigenkapital position (OR Art. 959a Abs. 2 Ziff. 3 lit. f and lit. g), which is what makes the statement foot before a year-end close has run. WHAT IT IS NOT: the statutory minimum structure. Twenty-four individual sub-positions are prescribed "einzeln und in der vorgegebenen Reihenfolge" (flüssige Mittel, Forderungen aus Lieferungen und Leistungen, aktive Rechnungsabgrenzungen, and so on) and are not modelled: only the seven groupings above are, with raw account lines under them. The two Absätze that prescribe those 24 are OR Art. 959a Abs. 1 and OR Art. 959a Abs. 2, and A08 reaches the grouping level of each and stops there. On the shipped Kontenrahmen KMU the account numbers happen to ascend in the statutory order, so the output looks right by coincidence; on a renamed or renumbered chart it names none of the required positions and may order them arbitrarily, and no reconciliation flag can see the difference because the statement still foots either way. So do not report this output as OR-conformant, as the minimum structure, or as ready to file. Pass compareTo={asOf:'YYYY-MM-DD'} for the prior-date column. This is the filed statement, so it accepts no alternate grouping.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | Yes | ||
| compareTo | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description richly discloses behavior: grouping level, positive lines, running result as Eigenkapital, foots before year-end close, and the limitation that on renamed charts the ordering may be arbitrary with no reconciliation flag to detect it. This goes far beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but dense, with each sentence carrying substantive detail. It is front-loaded with the purpose and structured with 'WHAT IT IS NOT' and a clear instruction for compareTo. While verbose, it avoids fluff and all content is relevant to correct usage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and lack of output schema, the description fully describes the output structure (sections, headings, account lines), the running result behavior, and the compareTo feature. It also covers limitations that could affect interpretation. Nothing essential is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. It explains the compareTo parameter's purpose (prior-date column) and implicitly clarifies asOf as the report date. WorkspaceId is self-evident. It adds meaning for compareTo, which is the least obvious parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool produces a balance sheet as of a date, grouped into specific OR Art. 959a categories with statutory headings and residual sections. It clearly distinguishes the output from the statutory minimum structure, making its purpose unambiguous and differentiating it from a generic report.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides context for when to use the tool: 'This is the filed statement' and instructs to pass compareTo for a prior-date column. It also explicitly cautions against reporting the output as OR-conformant, which is a clear when-not-to-use. However, it does not explicitly compare to sibling tools like trial_balance or income_statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bank_channel_connectA
Open or advance a bank channel, EBICS by default or the managed bLink rail with channelKind='managed_blink' (A37). P8-gated: pass confirm=true or enable the approval dial. EBICS (the SMPG 6.1 key ceremony, keyed to the bank CONTRACT host/partner/user, never an account): from nothing generate three RSA key pairs locally, create the connection, route the given A19 accounts, and file the INI letter as a document to sign and post to the bank (state keys_generated; with a transport wired it sends INI+HIA and lands pending_bank_activation); from pending_bank_activation run HPB and, with confirmBankKeys=true and a hash match, flip to active (a mismatch is bank_keys_mismatch, a hard stop). Managed (channelKind='managed_blink', keyed to the bank via bankRef with scopes ['ais'] or ['ais','pss']): with the owner-gated cloud tier ON it asks the relay for a fresh consent URL and lands consent_pending, and the next poll flips to active once the customer grants consent in their own e-banking; with the tier OFF it returns the OP4 shape {ok:false,error:'cloud_tier'} with no side effects. TILL never sees, asks for, or stores a bank credential on either rail; never returns key material; posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| host | No | ||
| scopes | No | ||
| bankRef | No | ||
| confirm | No | ||
| keyLength | No | ||
| channelKind | No | ||
| workspaceId | Yes | ||
| connectionId | No | ||
| idempotencyKey | Yes | ||
| confirmBankKeys | No | ||
| routeBankAccountIds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses the state machine (keys_generated, pending_bank_activation, consent_pending, bank_keys_mismatch hard stop), side effects (files INI letter, sends INI+HIA when wired, posts nothing), and security invariants (never stores credentials or returns key material). It also documents the cloud_tier error shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and runs as one long paragraph, but almost every clause adds program-state or security information needed to call the tool correctly. It is front-loaded with the core action and then structures the two rails in order; a light bulleted split would improve readability without changing substance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a stateful tool with 11 parameters, no annotations and no output schema, the description is remarkably complete: it names defaults, required flags, activation conditions, a hard-stop error, an error response shape, and explicit no-side-effect cases. An agent has enough context to decide whether to call it and what to pass.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the full burden, and it delivers: channelKind selection, confirm and confirmBankKeys flags, bankRef/scopes for managed mode, host/partner/user contract keying, and account routing are all explained with behavioral meaning. This goes well beyond the bare schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb and resource ('Open or advance a bank channel') and immediately defines scope by naming the two rails: EBICS by default and managed bLink via channelKind='managed_blink'. This is concrete enough to distinguish the tool from bank_channel_status/disconnect/directory without relying on the title.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear conditional context: EBICS keyed to contract host/partner/user, managed rail keyed to bankRef with scopes, P8-gated via confirm=true, and cloud-tier ON/OFF determining whether it asks for a consent URL or returns cloud_tier. It does not explicitly name sibling alternatives for when to check status or disconnect, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bank_channel_directoryARead-only
Look up a Swiss bank in TILL's static, in-package EBICS directory for the connect wizard (A36): substring or BIC match over the curated set, returning each bank's names, BIC, the marketing name of its EBICS channel, the segments it offers EBICS to, where to post the signed INI letter, known protocol quirks, and an EXPLICITLY UNVERIFIED fee note (verified:false always). The per-contract host URL and Host ID are null in v1 (the bank issues them on your signed EBICS contract; enter them in the manual fields). Reads static data only: opens NO socket, needs no idempotency key, is never a gate on the ceremony. An empty query returns the whole set; no match returns an empty array, not an error.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses extensive behavioral traits: it 'opens NO socket,' 'needs no idempotency key,' is never a ceremony gate, always returns verified:false on fee notes, returns null host URL/Host ID in v1, and defines edge-case behavior (empty query returns whole set; no match returns empty array, not an error). This far exceeds what the annotation provides and contains no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause carries distinct, decision-relevant information—no filler. The core lookup purpose is front-loaded, followed by return-content details, data caveats, behavioral guarantees, and edge cases in logical order.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema and 0% parameter coverage, the description fully covers return content (names, BIC, channel name, segments, INI letter posting, quirks, fee note), edge-case behavior, network behavior, and where this fits in the EBICS ceremony. An agent has everything needed to invoke it correctly; only workspaceId semantics are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates substantially for the query parameter by defining its matching semantics ('substring or BIC match') and empty-query behavior. The required workspaceId parameter is not explained at all, which is a minor gap given it is the one required field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Look up a Swiss bank in TILL's static, in-package EBICS directory' with a defined matching mode ('substring or BIC match'). It also distinguishes itself from sibling connection tools by framing this as a pre-ceremony lookup for the connect wizard (A36) and explicitly noting it 'is never a gate on the ceremony,' separating it from bank_channel_connect, bank_channel_status, and bank_sync.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this is for the connect wizard (A36), reads static in-package data, and informs manual entry of host URL and Host ID once a signed contract exists. However, it stops short of explicitly naming alternatives or stating when-not-to-use this tool versus sibling tools like bank_channel_connect or bank_channel_status.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bank_channel_disconnectA
Retire or block an EBICS bank channel (P8-gated: pass confirm=true or enable the approval dial). retire is the local, terminal close: the channel flips to retired, its order history stays readable, and its keys are destroyed. block is the emergency stop: it issues the SPR administrative order and blocks the EBICS channel only (the file path stays open); re-initialisation via bank_channel_connect is the way back. Both idempotent. Posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | ||
| confirm | No | ||
| workspaceId | Yes | ||
| connectionId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden and does so thoroughly. It discloses the P8 gating requirement (confirm=true or approval dial), the consequences of retire (keys destroyed, history stays readable), the side effects of block (SPR order issued, EBICS channel blocked while file path stays open), idempotency, and that it posts nothing. This is exemplary behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences pack an unusual amount of useful information: purpose, mode distinctions, consequences, gating, recovery, idempotency, and post behavior. It's dense but every clause earns its place; only a slight risk of overload from the parentheticals keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters, no enums, and no output schema, the description is remarkably complete: it covers gating, side effects, mode semantics, recovery, and idempotency. It does not describe the response shape, but the 'Posts nothing' note implies minimal output, and the behavioral coverage outweighs that omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds real meaning to mode (retire vs block) and confirm (P8 gating), but leaves workspaceId, connectionId, and idempotencyKey to be inferred from their names. The most critical parameter semantics are covered, but the gap for the remaining three params prevents a higher score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource pair: 'Retire or block an EBICS bank channel.' It immediately distinguishes between the two modes (retire vs block), which makes the tool's purpose unambiguous and separates it from related siblings like bank_channel_connect and bank_channel_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when each mode is appropriate: retire is the 'local, terminal close,' while block is the 'emergency stop.' It also names the recovery path via bank_channel_connect. It doesn't explicitly enumerate when to choose this tool over siblings, but the mode-level guidance is strong enough for an agent to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bank_channel_statusARead-only
Read EBICS channel health: per channel the connection state, last sync, any pending INI letter, uploads awaiting bank release, unmatched fetched statements, in-doubt uploads, recent bank_rejected orders, and the recent order log. Optionally scoped to one bank account, or filtered by order status (a saved view over the ebics_order log applies its stored status filter; an explicit status wins). A blocked channel is labelled as blocking only the EBICS channel, the file path still open. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| bankAccountId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses important behavioral details: the saved-view vs explicit-status precedence, the scoping semantics, and the nuanced caveat that a blocked channel only blocks EBICS while the file path remains open. These are non-obvious behaviors an agent needs to interpret results correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently packed, starting with the core purpose and then listing returned data in a readable enumeration. The scoping and filter rules are concise, and the final blocking caveat is valuable. Minor redundancy: 'Read-only' repeats the annotation, and the list could be tightened slightly without loss.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description appropriately lists the return content (connection state, sync, pending items, order log, etc.) and covers all optional parameter behaviors. It also includes a safety-relevant edge case (blocking semantics). Any missing details like pagination are not critical for a status read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates well for the optional parameters: bankAccountId is explained as scoping, and status/savedViewId get a full precedence rule. However, workspaceId is not mentioned (though it's obvious as a required workspace context), and the precise allowed values for status are not given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Read') and a clear resource ('EBICS channel health'), then enumerates exactly what health dimensions are returned (connection state, last sync, pending letters, uploads, etc.). This fully distinguishes it from sibling tools like bank_sync or bank_channel_connect, which are action-oriented, without requiring the reader to infer meaning.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains optional scoping via bankAccountId and the status/savedViewId filter precedence, but it never explicitly says when to choose this tool over alternatives like bank_channel_directory or list_bank_statements. The usage context is implied by the name and read-only nature, but no exclusions or alternative routing is stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bank_syncB
Pull pending camt.053/054 statements (plus the pain.002 status report and the HAC customer protocol) over an active EBICS channel and hand each statement to A20's import BYTE-FOR-BYTE, routed to the right A19 account by its IBAN. Persists every fetched file as a document before acknowledging the bank (crash-durable), and re-sync is A20's dedupe no-op. A statement whose IBAN matches no routed account is surfaced as unmatched_account, never dropped. With no transport wired it reports needs_bank_transport and the file path stands. Writes statements only through A20; posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| dateRange | No | ||
| workspaceId | Yes | ||
| connectionId | No | ||
| bankAccountId | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses side effects: it persists documents before acknowledging the bank (crash-durable), re-sync is a dedupe no-op, unmatched accounts are surfaced as unmatched_account, and it writes only through A20. This exceeds the typical disclosure and gives an agent confidence about safety and idempotency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph with multiple clauses and acronyms (A20, A19, camt, pain, HAC) that may confuse. It front-loads the primary action but becomes convoluted. It could be split into clearer sentences without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 5 parameters and no output schema, the description leaves out essential invocation details: what each parameter means, what the response contains, and what 'transport' refers to. It covers process and error cases but not how to actually call it. An agent would struggle to pick the right account or connection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does not explain workspaceId, dateRange, connectionId, or bankAccountId. The only indirect reference is to idempotencyKey via the dedupe no-op mention, but it does not clarify format or purpose. This is a major gap for correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: pulling pending camt.053/054 statements plus pain.002 and HAC protocol, and importing them to A20. It specifies the resource types and the routing to A19 accounts, which is specific and distinguishes it from generic 'list' or 'import' siblings, though it does not explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives like import_camt or list_bank_statements. It implies a sync/import use case but does not provide exclusions or conditions. The only contextual hint is about re-sync being a no-op, which is about idempotency, not tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billing_generate_invoiceA
Rechnungsentwurf aus Zeit erstellen (B02): turns the selected approved+billable+unbilled entries (one contact, one currency) into an A11 invoice DRAFT and flips them approved to billed with invoice_line_id set, in one transaction. Delegates to A10 createDocument (P3: no journal entry, no VAT amount, no total minted here; A11 to A02 post at issue). groupBy (entry|phase|project|day) shapes the lines; each line is qty 1 with the group value as its price, so preview equals line equals WIP to the Rappen. Refuses empty_selection, mixed_contacts, currency_mismatch, already_billed (strict, no partial invoice); idempotent on idempotencyKey. Always stops at a draft (P8); gated on A24 billing.generate.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| groupBy | No | ||
| contactId | Yes | ||
| throughDate | No | ||
| workspaceId | Yes | ||
| timeEntryIds | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and discloses mutation (approved→billed with invoice_line_id), transaction atomicity, idempotency on idempotencyKey, validation refusals, draft-only outcome, and permission gate. This is rich behavioral disclosure beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-rich but not concise; it bundles many domain codes (A10, A11, P3, P8, A02, A24, B02, WIP, Rappen) into a single unbroken paragraph. Every sentence adds unique behavioral details, but structure could be improved for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex billing mutation with no output schema and no annotations, this description is comprehensive about state transitions, validation, grouping, and permissions. It omits the return value/response contract and does not clarify the throughDate parameter, leaving some gaps for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains groupBy options (entry|phase|project|day), idempotencyKey semantics, contactId scope ('one contact'), and timeEntryIds scope ('approved+billable+unbilled entries'). It leaves workspaceId, actor, and throughDate undescribed, though these are less central.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific, multi-step behavior: turns approved+billable+unbilled time entries into an A11 invoice draft and flips their status to billed. The B02 code and 'Always stops at a draft (P8)' clearly differentiate it from posting/issuing tools in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: it creates a draft, delegates to A10 createDocument, and explicitly notes what it does not do (no journal entry, no VAT amount, no total minted), with posting deferred to A11→A02 at issue. It also states the gating permission (A24 billing.generate) and strict refusal conditions, but does not explicitly name an alternative tool to use instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billing_release_timeA
Zeit aus einem Entwurf freigeben (B02): reverts the entries backing a DRAFT invoice (by invoiceId or lineIds) from billed to approved and clears invoice_line_id, returning them to the unbilled pile and WIP. Called by A11 cancel and directly. Refuses a finalised invoice with invoice_not_draft (corrections there are A11 credit notes); releasing already-released lines is a no-op (idempotent). Gated on A24 billing.generate.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| lineIds | No | ||
| invoiceId | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the mutation (reverts billed to approved, clears invoice_line_id), the idempotent no-op behavior, the refusal condition (invoice_not_draft), and the permission gate (gated on A24 billing.generate). This is rich behavioral disclosure beyond simple 'releases draft time'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and packs a lot of information into one long sentence. It is efficient, but the German prefix 'Zeit aus einem Entwurf freigeben (B02)' may confuse non-German users, and the internal references (A11, A24) are unexplained. Despite this, every clause adds value and there is no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description covers the core behavior, error condition, and permission requirement. However, it does not explain the return value, whether the draft invoice header remains after release, or the relationship between invoiceId and lineIds when both are supplied. It is complete enough for most calls but leaves edge-case ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the two primary selection parameters, invoiceId and lineIds, which is the core semantic choice. However, it does not explain workspaceId, idempotencyKey, or actor, and does not clarify whether invoiceId and lineIds are mutually exclusive or what happens if both/neither are provided. It adds meaning but leaves some gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: reverts time entries backing a draft invoice from billed to approved and clears invoice_line_id. It names the explicit resource (DRAFT invoice, entries) and distinguishes itself from alternative correction paths (A11 credit notes for finalised invoices). An agent can tell this tool releases draft time entries rather than generating invoices or performing credit note corrections.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says it is 'Called by A11 cancel and directly' and gives a clear exclusion: for finalised invoices, corrections belong to A11 credit notes. It also notes idempotency, so an agent knows repeated calls are safe. This is direct, explicit when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billing_unbilled_previewARead-only
Unverrechnete freigegebene Zeit anzeigen (B02): the approved, billable, not-yet-billed time_entry pile (B01), grouped contact then project then phase, each entry valued round-once (P2) from its snapshotted rate and every subtotal an integer sum. Filter by contactId, projectId, a throughDate cut on started_at, or a groupBy hint. An empty pile is { groups: [], totalRappen: 0 }, a healthy state. B02 invoicing input; gated on the A24 billing.read capability.
| Name | Required | Description | Default |
|---|---|---|---|
| groupBy | No | ||
| contactId | No | ||
| projectId | No | ||
| throughDate | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes non-mutation, and the description adds meaningful behavior beyond that: empty-result shape, grouping order, snapshot-rate valuation, integer subtotals, and an authentication gate. It does not disclose pagination or non-empty entry fields in detail, but the added context is substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the core purpose, but it is a long run-on mixing German and English and relies on opaque internal codes (B01, B02, P2, A24). Every clause adds information, yet the structure hurts readability and could be clearer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description covers purpose, grouping, valuation, filtering, empty response, and permission gate, which is a strong amount of context. Missing concrete groupBy values, date format, and non-empty response details mean an agent still has to infer part of the contract.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage and no enums, the description does add semantics for contactId, projectId, throughDate, and groupBy, including that throughDate cuts on started_at. However, workspaceId is undocumented, groupBy allowed values are left as a vague 'hint', and date format / filter combination behavior is not specified, so it only partially compensates for the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete operation ('anzeigen' / preview), identifies the resource ('approved, billable, not-yet-billed time_entry pile'), and specifies the grouping order. The phrase 'B02 invoicing input' plus the 'not-yet-billed' qualifier clearly separates it from related billing actions like billing_release_time or billing_generate_invoice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: it is the read-only preview feeding B02 invoicing, gated on the A24 billing.read capability, and is intended for approved, billable, unbilled time. It does not explicitly name alternatives or exclusion conditions such as 'use billing_wip_report for WIP', so it stops short of full when/not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
billing_wip_reportBRead-only
WIP-Bericht der angefangenen Arbeiten (B02): approved-but-unbilled value as of a date, per project and its client, purely (P5). wipRappen is the integer sum of round-once entry values, plus minutes and the oldest entry age. Informational only: B02 never posts the OR 960c entry, a Treuhänder posts it manually from these numbers. Gated on A24 billing.read.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| contactId | No | ||
| projectId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds that it never posts entries and that a Treuhänder posts manually, reinforcing the read-only nature. It also explains the wipRappen calculation (integer sum of round-once entry values, plus minutes and oldest entry age). No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise at two sentences and front-loads the primary purpose. However, it is dense with domain jargon (B02, P5, OR 960c, wipRappen) that may obscure clarity for an agent unfamiliar with Swiss accounting terminology.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
It covers the report's purpose, informational nature, gating, and calculation. However, it omits parameter semantics and the output format. Since there is no output schema, the description should provide more detail on the response structure, but it does not.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain parameters like asOf, contactId, projectId, or workspaceId. It hints at 'as of a date' and 'per project and client', but does not map these to the schema fields. Given the lack of schema descriptions, the description should compensate but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reports approved-but-unbilled WIP value as of a date, per project and client. It names a specific resource (WIP) and verb (report). However, it does not explicitly differentiate from the sibling billing_unbilled_preview, which likely serves a similar purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides useful context: 'Informational only: B02 never posts the OR 960c entry' and mentions gating on A24 billing.read. However, it does not explicitly state when to use this tool versus alternatives like billing_unbilled_preview or when not to use it. The guidance is implied but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bootstrap_workspaceC
One-call agent setup: parse a short description into a workspace and VAT config.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| description | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, but it only says it 'parse[s]' a description into workspace and VAT config. It does not disclose whether the workspace is created, whether existing configuration is overwritten, whether the operation is destructive, or how the idempotencyKey affects behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or repetition. It is appropriately compact, though its brevity contributes to the lack of behavioral and parameter detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a high-level setup tool with no output schema or annotations, yet the description omits essential operational details: what is created or modified, whether the call is idempotent, what the response contains, and what configuration the agent can expect afterward. It is not complete enough for an agent to invoke safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only hints at the 'description' parameter via 'short description.' It provides no guidance on the required idempotencyKey or the optional name parameter, including their formats or roles in the bootstrap process.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: it 'bootstrap[s]' a workspace and VAT config from a short description. This is meaningfully clearer than a tautology, though it does not explicitly differentiate from closely related siblings like create_workspace, vat_configure, or set_vat_config.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'One-call agent setup' implies this tool is intended to replace a multi-step workspace-plus-VAT configuration flow, but it never names alternatives or states when not to use it. Usage context is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capture_commitA
Commit a reviewed capture into a DRAFT: target.kind='vendor_bill' delegates to A17 create_vendor_bill (+ attach_receipt + the E00 entity link that derives OR 958f retention), and target.kind='expense_line' delegates to E02 (onto target.claimId, or a fresh draft claim for target.employeeId). Requires the delegated verb's own capability (post for a bill, spesen.submit for an expense), re-checked inside the delegation. Posts NOTHING: the target lands as a draft and posting it is A17's separately-dialled step (P8). corrections (each {key,value}) land as operator fields before delegation. Idempotent: a double-commit returns the original target_id and never creates a second draft.
| Name | Required | Description | Default |
|---|---|---|---|
| target | Yes | ||
| captureId | Yes | ||
| corrections | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and discloses critical behaviors: it posts nothing, lands the target as a draft, is idempotent, re-checks the delegated verb's capability, and applies corrections as operator fields. This is substantial transparency. It doesn't describe failure modes or invalid target.kind handling, but the key operational behavior is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense yet information-rich, packing many behavioral details into three sentences. It front-loads the core commit action and delegation logic. The heavy use of parentheticals and punctuation makes parsing harder, but every sentence contributes value. A more structured layout would earn a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 5 parameters, nested targets, and cross-tool delegation, the description covers the essential context: delegation paths, draft behavior, permission requirements, idempotency, and correction placement. Gaps include no output schema details, no failure-mode guidance, and incomplete description of captureId/workspaceId roles. These are moderate but not severe for a tool in a well-defined pipeline.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains target semantics (the two kinds and their delegation), corrections as {key,value} operator fields, and the idempotencyKey behavior. However, captureId and workspaceId are not elaborated, and the nested target structure is only partially described. The compensation is strong but not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb (commit) and resource (a reviewed capture into a DRAFT), and explains the two delegation paths based on target.kind. It distinguishes the tool from capture_document, capture_extract, and capture_discard siblings, but doesn't explicitly name the alternatives, so the differentiation is implicit rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use this tool: committing a reviewed capture, with target.kind selecting the delegated verb. It also clarifies the tool's non-posting behavior and idempotency, but it doesn't explicitly state when not to use it (e.g., use capture_discard to reject, post_vendor_bill to post). The delegation conditions are explicit, but exclusion guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capture_discardA
Discard a capture that is not a bookable document: flips a needs_review capture to discarded (terminal). The E00 document survives untouched (its retention is E00 law). Discarding a committed capture is refused with invalid_state (correct it on the A17/E02 side). A discarded capture is recoverable by re-uploading the same bytes (the dedupe key is scoped to non-discarded captures).
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| captureId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden for behavioral disclosure. It discloses the terminal nature of the state, side effects ('The E00 document survives untouched'), error behavior ('refused with invalid_state'), and recovery semantics ('recoverable by re-uploading the same bytes' with dedupe scoping). This is exceptionally transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four dense sentences, each earning its place. The core purpose is front-loaded in the first sentence, followed by necessary side-effect, error, and recovery details. No fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the state transition, preconditions, error handling, side effects on the E00 document, and recovery path. For a state-changing operation with no output schema or annotations, it provides an agent with virtually everything needed to decide when and how to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. It does not explain 'reason' or 'idempotencyKey' directly, nor does it clarify how workspaceId/captureId relate to the state machine. The only indirect hint is 'dedupe key is scoped to non-discarded captures', which touches on idempotency but not explicitly tied to a parameter. This leaves significant gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb and resource ('Discard a capture') and clearly defines the state transition: 'flips a needs_review capture to discarded (terminal)'. It also distinguishes from bookable documents and references the E00 document surviving untouched, separating this from commit-like tools. The purpose is unambiguous and self-contained.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage conditions: use for captures that are 'not a bookable document' and in 'needs_review' state. It explicitly states when not to use it: 'Discarding a committed capture is refused with invalid_state' and offers an alternative path ('correct it on the A17/E02 side'). This is clear when/when-not/alternatives guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capture_documentA
Ingest a supplier document into the Belegeingang queue (Belegerfassung). contentBase64 is the raw bytes; mime must be one of application/pdf, image/jpeg, image/png, image/heic, image/tiff (the queue refuses anything else, unlike E00). The file is stored via E00 and the deterministic pass decodes any Swiss QR Code payload present as text (creditor, IBAN, amount, currency, QRR/SCOR reference revalidated by check digit, and the Swico S1 billing information), landing each value as a provenance-stamped field. A byte-identical re-drop returns the existing capture ({duplicate:true}) and writes nothing; a hash matching only a discarded capture creates a fresh one cross-linking it (recovery). Posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| mime | Yes | ||
| filename | No | ||
| workspaceId | Yes | ||
| contentBase64 | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden; it discharges it by disclosing storage via E00, QR-code decoding and provenance stamping, duplicate-vs-recovery behavior, queue MIME rejection, and the fact that it posts nothing. This is unusually rich and goes well beyond what the schema provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence adds either parameter constraints, storage/extraction behavior, or duplicate/recovery semantics, and the main purpose is front-loaded. It is long only because the tool is complex; no filler is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, this description covers purpose, accepted MIME types, storage, QR-extraction side effects, duplicate handling, and recovery. It leaves gaps around the exact meaning of workspaceId/idempotencyKey and the success return payload, but overall an agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds strong semantics for contentBase64 ('raw bytes') and mime (explicit allowed list), which is essential since schema coverage is 0%. However, workspaceId, filename, and idempotencyKey are not named or explained, and the idempotency behavior is described only as 'byte-identical re-drop' rather than tied to the required idempotencyKey parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific action ('Ingest a supplier document into the Belegeingang queue') with a concrete resource and destination. It also contrasts with E00 and ends with 'Posts nothing,' which helps an agent distinguish it from posting/commit tools. This is specific enough to separate it from sibling tools like capture_extract or capture_commit without opening their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear context for when to call it (ingesting a supplier document into the capture queue) and an explicit MIME restriction, but it does not name alternatives such as capture_extract or capture_commit nor say when those should be used instead. 'Unlike E00' references an unknown external system, not a sibling MCP tool, so routing guidance is mostly implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
capture_extractA
Re-run or augment a capture's extraction. source='qr' re-runs the deterministic pass; source='agent' takes caller-supplied fields (each {key,value,confidence}, validated against the CAPTURE_FIELD_KEY registry); source='operator' takes review-pane corrections (each {key,value}, landing as operator/high); source='local_model' degrades honestly to needs_local_runtime (no on-device model is wired in the core). Merge rules are fixed: an operator value is never overwritten by a machine source, higher confidence wins, and a deterministic source outranks a probabilistic one at equal confidence; the loser is kept as a superseded row (auditable disagreement, never a silent drop). Re-running on unchanged input is a no-op.
| Name | Required | Description | Default |
|---|---|---|---|
| fields | No | ||
| source | Yes | ||
| captureId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers: it details merge rules (operator never overwritten, higher confidence wins, deterministic over probabilistic), explains that losing values are kept as superseded rows (never silently dropped), and states idempotency (no-op on unchanged input). It also discloses that local_model degrades honestly to needs_local_runtime. This is exemplary behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but each sentence adds value, and the main action is front-loaded. It could be broken into bullet points for readability, but it avoids fluff and is appropriately sized for the complexity described.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of output schema and annotations, the description covers behavior, source semantics, merge rules, and idempotency comprehensively. It would be improved by stating the return value (e.g., the updated capture or extraction result) or error conditions, but the core calling context is well covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters, and it does for 'source' and 'fields': it defines each source value and the exact field structure (key/value/confidence for agent, key/value for operator). It also mentions validation against a registry. However, workspaceId, captureId, and idempotencyKey are not explicitly explained beyond their names and the no-op hint, leaving some room for inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific operation (re-run or augment) on a resource (capture's extraction) and enumerates the four source modes with distinct behaviors. It clearly distinguishes from sibling capture tools (capture_document, capture_commit) by focusing on extraction augmentation rather than initial creation or commit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly defines when to use this tool (when you have an existing capture and want to add or re-run extraction) and explains each source type's purpose. It lacks explicit 'do not use when...' statements or named alternatives, but the context is clear enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
checklist_abandonA
Abandon a checklist run with a reason (a wrong period was started). The run stays listed under the abandoned filter and contributes nothing to the attention hub; it is never deleted, and every write on it afterwards refuses run_abandoned. Start the right period with checklist_start. Gates on manage_checklists. CONSEQUENCE: Abandons the run with your reason; it stays listed and stops feeding the attention hub.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| reason | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it excels: it discloses that the run stays listed under the abandoned filter, contributes nothing to the attention hub, is never deleted, and that every subsequent write refuses run_abandoned. These consequences are explicit and comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately long but every sentence earns its place: it front-loads the purpose, then explains consequences, the alternative, and the permission. It is well-structured and not redundant, though slightly verbose with the CONSEQUENCE recap.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the purpose, alternative, permission, and post-abandonment behavior. It does not explicitly state prerequisites (e.g., the run must be in progress), but the context 'a wrong period was started' implies it. Without an output schema, return values are not required, so it is reasonably complete for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all parameters. It explains the reason parameter ('with a reason') but does not elaborate on workspaceId, runId, or idempotencyKey beyond their names. While these are common, the description does not clarify their meaning or usage, leaving the agent to infer them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (abandon) and the resource (a checklist run), and adds the specific context (a wrong period was started). It distinguishes from checklist_start by naming the alternative directly. It also clarifies the run's fate (stays listed, no attention hub contribution, never deleted), leaving no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly instructs to use checklist_start to start the right period, and states the permission gate (manage_checklists). It implies the correct use case (when a wrong period was started) and, by contrast, when not to use it. This provides clear routing to alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
checklist_getARead-only
Read one checklist run with every item derived LIVE: a system check item is done while its check passes (drafts, bank reconciliation, tax codes, the vat_filed lock, the month or year lock), a verb or preview item is done while the hash the engine bound still equals the live read (stale otherwise), a posting item is done while its probe finds a live, unreversed artefact in the ledger, a validation item is done on pass (a warn needs a live acknowledgement), a choice item is done while a human answered or the books derive the answer, a sign-off item is done while a live sign-off stands, and an item whose governing choice carries another answer reads excluded. Returns the items in journey order, the next actionable item (nextItemId), the counts, the anchor hash and the derived run status (open | done | abandoned). Gates on read_books.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description adds extensive behavioral detail: how each item type is derived (system checks, verb/preview, posting, validation, choice, sign-off), the order of returned items, and the gate on read_books permission. This reveals side-effect-free behavior and computational logic that annotations do not cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a dense paragraph but not overly long given the complexity. It front-loads the main purpose ('Read one checklist run with every item derived LIVE') and then details item derivation and outputs. It is concise enough but could benefit from bullet points for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the description is thorough: it explains all item types, return fields, and the permission gate. It lacks parameter semantics (covered separately) but otherwise covers what an agent needs to understand behavior and outputs. Very complete for a read operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, and the description never explains the two parameters (workspaceId, runId). It does not clarify that workspaceId identifies the workspace and runId identifies the checklist run. Since schema coverage is low and the description must compensate, this is a notable gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states it reads one checklist run and derives every item live, then lists the output fields (items, nextItemId, counts, anchor hash, derived status). This is a specific verb+resource that clearly distinguishes it from sibling tools like checklist_list (list runs) and checklist_start (start runs).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly conveys what the tool does, so an agent can infer when to use it (to fetch the current derived state of a single run). However, it does not explicitly state when not to use it or mention alternatives, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
checklist_item_completeA
Complete one checklist item. A verb or preview item is completed by the ENGINE re-running the read and binding its hash as evidence; a caller-supplied evidence.ref must agree (evidence_mismatch otherwise); a preview whose read refuses answers read_refused. A choice item stores the answer given as evidence {kind: choice, ref: } (invalid_choice names the options; choice_locked when a governed posting or sign-off is already done). A warn validation records the acknowledgement as evidence {kind: signoff, ref: } bound to the figures (acknowledge_needs_reason without one); a block validation refuses check_not_passed. A sign-off item records an append-only sign-off: abstimmung_reviewed needs the live bridge check to pass (check_not_passed) and binds the return hash, statements_signoff binds the statements hash, settlement_booked needs evidence.ref naming the bank transaction or entry (evidence_required), eportal_filed needs evidence {kind: filed_attestation, ref: YYYY-MM-DD} not before the export (attestation_before_export) and a gv_attestation needs evidence {kind: gv_attestation, ref: YYYY-MM-DD} (a date before the statements sign-off needs evidence.reason); both refuse a second live attestation (already_attested). A system check or a posting item cannot be completed by hand (check_item_live: the domain verb is the act); an excluded item refuses item_excluded. Refuses prerequisite_open naming the blocking item and run_abandoned on an abandoned run. Gates on manage_checklists. CONSEQUENCE: Records the item as done under your name, binding the evidence the engine computed or the sign-off you stand behind; the ePortal attestation is the statutory filing claim and drafts for the owner under the vat-file dial.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| itemId | Yes | ||
| evidence | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden, and it delivers: it discloses append-only sign-offs, evidence binding under the caller's name, statutory filing implications, prerequisite/live-bridge checks, and a full set of failure modes. This far exceeds a generic 'completes an item' statement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence is a crisp summary and the rest is dense, high-signal detail without filler. It is a single long paragraph rather than structured bullets, which slightly reduces scannability, but every clause earns its place by describing conditions, errors, or consequences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with no annotations and no output schema, the description is unusually complete: it covers every item category, required evidence shape, failure codes, permission gate, and side-effect consequence. An agent can determine both whether to call it and what evidence to construct.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only provides types, so the description compensates by defining evidence.kind/ref/reason semantics, date formats, optionId references, and bank-transaction refs. It does not explicitly restate workspaceId/runId/itemId/idempotencyKey, but these are conventional identifiers and the critical parameter (evidence) is thoroughly specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Complete one checklist item') and then differentiates the action by item type, from verb/preview/choice/warn/sign-off to system check/posting. This makes it unmistakable from sibling checklist tools like checklist_item_skip or checklist_item_reopen.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly scopes when completion applies and explicitly states when it cannot be used: system checks and posting items cannot be completed by hand, excluded items refuse, and an abandoned run is refused. It does not explicitly name checklist_item_skip/reopen as alternatives, but the conditions and error cases effectively tell an agent when not to invoke this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
checklist_item_reopenA
Return a done or skipped checklist item to open. A live sign-off on it (the bridge review, the ePortal attestation, the payment, an acknowledgement) is voided with reason reopened and kept as history; a bound verb or preview evidence is cleared; a human choice answer is cleared back to the derived one (choice_locked while a governed posting or sign-off is done). A system check or posting item cannot be reopened by hand (check_item_live): it flips with its check or probe. Gates on manage_checklists. CONSEQUENCE: Reopens the item and voids its live sign-off; the earlier sign-off stays on record as voided.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| itemId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and delivers richly: it discloses that live sign-offs are voided and kept as history, bound verb/preview evidence is cleared, human choice answers revert to derived ones, and the item reopens with the earlier sign-off remaining as voided. It even explains the special behavior for system items. This is far beyond typical descriptions and gives the agent a clear model of side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but each sentence adds essential behavioral information. It is logically organized: primary purpose, side effects, exclusions, permission gate, and consequence. Some domain jargon (ePortal attestation, choice_locked, check_item_live) increases density but is appropriate given the complexity. It could be tightened with bullet points but remains well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the absence of annotations, and no output schema, this description covers all critical aspects an agent needs: state transition, side effects on sign-offs and evidence, handling of human vs. derived answers, restrictions on system items, permission requirement, and the final consequence. Nothing crucial for a correct call is left unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description does not explain any of the four parameters (workspaceId, runId, itemId, idempotencyKey). The parameter names are somewhat self-explanatory in context, but the description adds no semantic detail about how they combine to target a checklist item or why idempotencyKey is needed. For a tool with zero schema documentation, this is a notable gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Return a done or skipped checklist item to open.' This clearly defines the state transition and differentiates it from sibling tools like checklist_item_complete, checklist_item_skip, and checklist_abandon. It is unambiguous about the action being taken.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states the scope of use ('done or skipped checklist item') and explicitly excludes system check or posting items ('cannot be reopened by hand'), giving an agent a clear boundary. It also mentions the required permission gate ('Gates on manage_checklists'). However, it does not explicitly name alternative tools or say 'use X instead' when this tool is inappropriate, which keeps it just shy of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
checklist_item_skipA
Mark one checklist item as not applicable, with a recorded reason (skip_needs_reason without one). The undeletable items (the computed return, the period lock, the payment, the statements sign-off, the seal) refuse (undeletable): the canon is waived consciously, never dropped. An excluded item refuses item_excluded. A live sign-off on the item is voided, never deleted. Gates on manage_checklists. CONSEQUENCE: Marks the item as not applicable with your reason; it stays visible in the run and in the audit log.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| itemId | Yes | ||
| reason | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden and addresses it thoroughly: it discloses that undeletable items refuse, excluded items refuse, live sign-offs are voided (not deleted), and the item remains visible in run and audit log. It also hints at a 'skip_needs_reason' config. This is rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but structured: it leads with the core action, then lists refusal conditions, then the permission gate, then the consequence. The parenthetical '(skip_needs_reason without one)' is cryptic and could be clearer, but overall it is economical and organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, this description covers the essential behavioral contract: what it does, what it cannot do, side effects (voiding sign-offs), and audit visibility. It doesn't specify return values or error handling beyond refusals, but those are secondary for the AI to decide invocation. It's sufficiently complete for decision-making.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explains that the tool operates on a single checklist item and records a reason, which adds meaning to the `itemId` and `reason` params. But it does not clarify `runId`, `workspaceId`, or `idempotencyKey` beyond what their names imply. With 0% schema coverage, more compensation was needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Mark one checklist item as not applicable'. It distinguishes from sibling tools by specifying the exact semantic ('not applicable') and by listing what it refuses to do (undeletable items, excluded items, etc.), making it clear this is not a delete or complete operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives context on when this can be used: it requires 'manage_checklists' gate and a recorded reason. It also clarifies exclusions (undeletable items, excluded items) so an agent knows when not to attempt. It doesn't explicitly name sibling alternatives like 'complete' or 'reopen', but the behavior is clearly distinct.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
checklist_listARead-only
List this workspace's checklist runs as metadata rows (template, period, derived status, next item, counts), open first, then done, then abandoned, newest period first. Optional templateId and status (open | done | abandoned) filters. Abandoned runs are listed, never deleted. Gates on read_books.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| templateId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds several behavioral traits beyond the readOnlyHint annotation: abandoned runs are listed but never deleted, ordering is specified (open, done, abandoned, newest period first), and access gates on read_books. These details are non-obvious and provide valuable operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each adding meaningful information: the listing purpose and output fields, the ordering, and the filters/behavior/permissions. No fluff or redundancy; it is efficiently front-loaded with the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the listing purpose, output fields, ordering, filters, retention behavior, and permission gate. It does not mention pagination or response limits, but for a list tool without an output schema, it provides a solid understanding of what is returned and how it behaves.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by explaining the templateId and status filters, including the allowed status values (open | done | abandoned) that are not present in the schema. WorkspaceId is implicitly covered by 'this workspace's', and while not every parameter is detailed, the key filters are well explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists checklist runs with specific output fields (template, period, derived status, next item, counts) and ordering rules. It distinguishes this from sibling tools like checklist_templates, checklist_get, and checklist_start by focusing on the list operation for runs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage through its definition: listing runs with optional templateId and status filters. It provides clear context on filtering and ordering, though it does not explicitly name alternative tools or exclusions. The scope is clear enough for an agent to choose this over related checklist tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
checklist_startA
Start a checklist run for a template and a period. period is OPTIONAL: omitted, the engine picks the last ENDED period of the template's kind (an A07 label such as 2026-Q2 for vat_period, YYYY-MM for month, the fiscal year YYYY for year), which is what the seeded daily auto-start rule relies on. Creates every item with its owner, prerequisites and due date (the ESTV filing and payment items at period end + 60 days, Art. 71 Abs. 1 and Art. 86 Abs. 1 MWSTG; on a year the Umsatzabstimmung at + 180 days, the Berichtigung at + 240 days and the GV six months on, Art. 699 Abs. 2 OR). Idempotent on workspace + template + period: a second start returns the existing run with created:false. Refuses needs_vat_config without an A05 configuration and period_not_filable for a vat_period label that is not one of the year's filing periods (naming them); period_not_ended while a month or year has not ended; year_already_closed on a sealed fiscal year; year_close_in_progress for the last fiscal month while a year_close run for that year exists (the year run covers it; an existing December run is still returned). Gates on manage_checklists. CONSEQUENCE: Creates the checklist run for the period with every item and its statutory due date; a run is abandoned, never deleted.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | ||
| templateId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses idempotent behavior ('a second start returns the existing run with created:false', 'a run is abandoned, never deleted), permission gating on manage_checklists, and the side-effect consequence of creating every item with owners, prerequisites, and due dates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place, front-loading the core action first and then layering default behavior, idempotency, refusals, and consequences. It is long and packed with legal citations and error strings, which makes it slightly harder to scan but highly informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is remarkably complete for a tool with no output schema: it covers default period selection, item creation, idempotency, all relevant refusal conditions, permission gates, and the eventual fate of the run. An agent has enough context to invoke the tool correctly and anticipate errors.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Despite 0% schema coverage, the description richly explains the period parameter: optionality, default behavior, period kinds and formats, and error conditions tied to period values. It also clarifies idempotency key semantics implicitly by tying idempotency to workspace + template + period, and the parameter names themselves are self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Start a checklist run for a template and a period.' It distinguishes this from sibling checklist operations such as checklist_list, checklist_get, checklist_item_complete, and checklist_abandon by making clear this is the creation/start action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong context: when period is omitted the engine picks the last ended period, idempotency semantics, and an exhaustive list of refusal conditions (e.g., 'period_not_filable for a vat_period label that is not one of the year's filing periods' and 'gates on manage_checklists'). However, it does not explicitly name alternative tools such as month_end_checklist or say when to choose them instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
checklist_templatesARead-only
List the shipped checklist templates (vat_period, the MWST-Periode from open books to a filed, locked, paid return; the close templates month_close and year_close as they land): id, kind, label, description, period kind (vat_period | month | year), anchor (vat_return | statements) and item count. Describes the software, not the workspace. Gates on read_books.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description adds value by disclosing the permission gate ('Gates on read_books') and the scope (software vs. workspace). It does not contradict annotations and provides additional behavioral context beyond the structured metadata.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence that efficiently packs template names, returned fields, and the permission gate without fluff. It is front-loaded with the core purpose, though it could be slightly reorganized for readability. No redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lists the returned fields (id, kind, label, description, period kind, anchor, item count), which partially compensates for the missing output schema. However, it does not explain the workspaceId parameter, mention potential pagination or ordering, or clarify edge cases (e.g., empty template list). Given the tool's simplicity, this is acceptable but not thorough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage for the single required parameter workspaceId, and the tool description does not explain its purpose or impact. The phrase 'Describes the software, not the workspace' could confuse an agent about whether workspaceId is even relevant. The description fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('List'), the resource ('shipped checklist templates'), and enumerates specific templates (vat_period, month_close, year_close) and the fields returned. It also distinguishes itself from workspace-specific checklist tools by noting 'Describes the software, not the workspace,' making the scope unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by contrasting with workspace-specific checklists ('Describes the software, not the workspace'), but it does not explicitly name alternative tools (e.g., checklist_list) or provide when-to-use vs. when-not-to-use guidance. An agent must infer that this is for cataloging available templates rather than inspecting active checklists.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
clear_diagnosticsB
Delete every recorded error detail. Reports already written are kept.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It does disclose the destructive action ('Delete every recorded error detail') and the carve-out that already-written reports are kept, which is genuinely useful. However, it does not address reversibility, permissions, or side effects beyond those two points.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, with the core action front-loaded and the preservation caveat immediately following it. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no annotations and no output schema, the description is thin: it omits parameter semantics, reversibility, permission expectations, and any mention of the workspace scoping implied by the required workspaceId.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the required workspaceId parameter, but it never mentions it. The agent gets no help on what workspaceId refers to or how it constrains which diagnostics get deleted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource ('Delete every recorded error detail') and adds a scoping note that reports are preserved. The sibling set includes get_diagnostics and set_diagnostics, and the destructive counterpart is implied by the verb, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives like get_diagnostics or set_diagnostics. There is no mention of prerequisites, recoverability, or situations where clearing diagnostics would be inappropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_monthA
Soft-close a month (reversible guardrail). CONSEQUENCE: Soft-closes the month: every later posting into it is refused until a human reopens it.
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly states the main effect: every later posting is refused until a human reopens it, and it is reversible. It does not mention idempotency behavior or permission requirements, but for a simple action it covers the key behavioral impact and potential side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no waste. The action is front-loaded, and the consequence is stated clearly. It is efficient and to the point, making it easy for an agent to grasp quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core behavior and reversibility, but it lacks parameter explanations and does not contrast with sibling tools like lock_period or reopen_month. There is no output schema, so the description should explain what the call returns, but it doesn't. It partially covers the context but misses critical details for a complete understanding.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain any of the three parameters (workspaceId, period, idempotencyKey). The agent must infer from parameter names alone, which is insufficient. The description adds no value for understanding what each parameter means or how to format them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Soft-close a month', and immediately clarifies the consequence (refuses later postings). It implicitly distinguishes from siblings like close_year and lock_period by emphasizing reversibility and the guardrail nature, making the tool's purpose clear and specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives context that this is a reversible guardrail and that later postings are refused, implying it's for month-end soft closing. However, it does not explicitly mention when to prefer this over lock_period, close_year, or other related tools, nor does it state any exclusions. The context is clear but lacks explicit guidance on alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
close_yearA
Hard-close a fiscal year: sweep the P&L into equity and seal the year. CONSEQUENCE: Sweeps the year's profit or loss into equity and seals the fiscal year; a sealed year cannot be reopened.
| Name | Required | Description | Default |
|---|---|---|---|
| year | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavior disclosure. It explicitly states a critical consequence: 'a sealed year cannot be reopened,' which warns of irreversibility. It also describes the action of sweeping P&L into equityUTEIR. This is strong for a closure tool, though it omits other effects like impact on locked periods or user permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise, two sentences, with the primary action up front and the critical consequence stated clearly. Every sentence adds essential information without redundancy or filler. Its structure is efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an irreversible year-end close operation, the description is incomplete. It lacks information on prerequisites (e.g., whether periods must be closed first), how idempotencyKey should be used, and what happens if the year is already closed. It also does not clarify the relationship to sibling tools like 'lock_period' or 'reopen_month', leaving the agent to guess the broader workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does not mention any of the three parameters (workspaceId, year, idempotencyKey) or their semantics. The parameter names are not explained, and the description adds no value beyond what the raw names suggest. This leaves the agent without guidance on required input format or purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action: 'Hard-close a fiscal year: sweep the P&L into equity and seal the year.' The verb+resource is explicit and the behavior is specific. It differentiates itself from sibling tools like 'close_month' by emphasizing 'fiscal year' and 'sealed', making the intent unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool (when closing a fiscal year) but does not provide explicit guidance on alternatives or exclusion criteria. For instance, it doesn't contrast with 'close_month' or 'prepare_period', nor does it mention prerequisites like completing month-end closes. The context is clear but not fully elaborated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
comment_entryA
Ask or note something about a posted journal entry without touching it: appends one comment to the entry review thread (Prüfvermerk) and never changes the entry itself or its review status. The Treuhänder queries, the client answers with a correcting reversal (reverse_entry), never an edit.
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| entryId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations available, the description carries the full disclosure burden and does so well: it appends exactly one comment, does not modify the entry, and does not change review status. It omits idempotency behavior despite requiring idempotencyKey, but the critical side-effect profile is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact: two sentences, front-loaded with the core behavior and the non-mutation guarantee. The German terms Prüfvermerk and Treuhänder add some opacity, but there is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple append-comment operation, the description covers side effects, non-destructive intent, and the correction workflow. However, because there is no output schema or annotations, the missing idempotency-key behavior, return value, and any constraints on comment text length leave the definition slightly incomplete for fully reliable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description must compensate, but it only loosely implies that text is the comment and entryId targets the posted entry. workspaceId and especially idempotencyKey semantics are not explained at all, leaving an agent to guess about scoping and duplicate prevention.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('appends one comment'), a specific resource ('entry review thread / Prüfvermerk'), and a clear non-goal ('never changes the entry itself or its review status'). This distinguishes it clearly from mutation and reversal tools such as reverse_entry.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear role-based guidance: the Treuhänder queries via this thread, and the client answers with reverse_entry rather than an edit. The description explicitly names reverse_entry as the alternative for correcting, which is the key routing decision. It does not exhaustively compare against every sibling, but the essential usage context is present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confirm_fieldA
Release a draft custom field so every screen and every caller can see it: the human half of draft-by-default. Confirming a field that is already live is a successful no-op.
| Name | Required | Description | Default |
|---|---|---|---|
| fieldDefId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It clearly states the primary side effect (visibility to all screens/callers) and the idempotency behavior (already-live confirm is a no-op). This covers the most critical behavioral traits, though it omits potential irreversibility or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. It front-loads the action and purpose, adds a memorable metaphor, and includes a key idempotency detail efficiently. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple state-transition tool with no output schema and no annotations, the description provides the core purpose, effect, and idempotency behavior. It lacks explicit parameter guidance and return-value details, but those are partly inferable from the schema and the tool's simplicity, making this nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention any of the three parameters (fieldDefId, workspaceId, idempotencyKey) or their semantics. While parameter names are somewhat self-explanatory, the description adds no meaning beyond the schema's field names and types, failing to compensate for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Release') and resource ('draft custom field'), and explains the effect ('every screen and every caller can see it'). It clearly distinguishes this from related field operations by framing it as the 'human half of draft-by-default,' which signals its role in the draft-to-live workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context that this tool moves a draft field to live status, and mentions that confirming an already-live field is a no-op. However, it does not explicitly name alternatives or state when not to use it, leaving the agent to infer the workflow from context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confirm_matchA
Settle an outgoing DEBIT against one or more vendor bills (vendorBillId, or allocations[] for a multi-bill split) through A14's record_payment, or annotate a LINK to an existing journal entry that already books the movement (entryId). A CREDIT-classified txn refuses with use_qr_queue: decide it in the A21 Abgleich queue (apply_qr_match / override_qr_match) instead. A settlement (vendorBillId/allocations) against a bank fact that is not a DBIT (e.g. a reversal-flagged returned payment, money arriving) refuses with wrong_direction; correct a returned payment via reverse_payment, or use entryId, which stays open to either direction. Idempotent per txn. CONSEQUENCE: Settles the bank debit against its vendor bills and posts the payment.
| Name | Required | Description | Default |
|---|---|---|---|
| entryId | No | ||
| bankTxnId | Yes | ||
| allocations | No | ||
| workspaceId | Yes | ||
| vendorBillId | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does well: it discloses refusal modes (use_qr_queue, wrong_direction), idempotency per txn, and the concrete consequence of settling and posting the payment. This is exactly the kind of behavioral context an agent needs beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded with the primary purpose, then covers exclusions and consequences. Every sentence adds information, though the heavy capitalization and domain-specific shorthand make it slightly harder to scan than necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and the absence of annotations, the description covers the main modes, error cases, idempotency, and consequences well. It is not fully complete because the allocation object's targetId/amountMinor semantics and the exact relationship to record_payment are left underspecified, but an agent can still handle common cases confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It explains the semantic difference between vendorBillId, allocations[], and entryId, and mentions idempotencyKey behavior. But it does not explain required workspaceId/bankTxnId or the allocation subfields targetId/amountMinor, leaving ambiguity for multi-bill splits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Settle an outgoing DEBIT against one or more vendor bills') and a clear alternative mode ('annotate a LINK to an existing journal entry'). It also distinguishes the tool from siblings like apply_qr_match and override_qr_match, so an agent can identify what confirm_match does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly routes CREDIT-classified transactions to apply_qr_match/override_qr_match and tells the agent to use reverse_payment or entryId when the bank fact is not a DEBIT. However, it does not directly clarify the relationship with the sibling record_payment, even though it says settlement happens 'through A14's record_payment'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_anonymiseA
Anonymise a contact on a valid revDSG deletion request, bounded by OR 958f: blanks the personal fields and redacts the activity bodies while keeping the row ids so posted-document FKs stay intact. Erases the whole merge identity (the contact plus every duplicate merged into it), so pass the SURVIVOR: a merge tombstone is refused and names it. Refuses while any unsettled receivable or obligation is still live (draft, issued, sent, accepted, confirmed, partially_paid), because an unpaid claim is an overriding interest and OR 958f Abs. 3 needs the retained record to stay readable.
| Name | Required | Description | Default |
|---|---|---|---|
| contactId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral disclosure burden. It discloses side effects, preservation of foreign keys, erasure of merge identities, and refusal conditions with legal justification. This is unusually transparent for a destructive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
All three sentences carry essential behavioral or usage information with no filler. The core action is front-loaded, followed by merge-identity details and refusal conditions packed efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive, legally bounded operation with no annotations and no output schema, the description covers preconditions, side effects, side effects, refusal cases, and the reason for refusal. It could additionally describe the success response or idempotency semantics, but the core invocation requirements are substantially addressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description compensates for the critical contactId parameter by requiring the survivor and rejecting merge tombstones. workspaceId and idempotencyKey are left to self-explanatory names, so the description does not fully cover all three parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the exact operation (anonymise a contact) and details what happens: personal fields are blanked, activity bodies are redacted, row IDs are kept, and the whole merge identity is erased. It also distinguishes itself from a merge tombstone by refusing the tombstone and requiring the survivor.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when to use it (valid revDSG deletion request bounded by OR 958f) and when it refuses (live receivables/obligations). It doesn't explicitly compare against sibling tools like contacts_merge, but it does tell the caller to pass the survivor rather than a merge tombstone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_importA
Import already-parsed contact rows with duplicate detection. Each row is validated like create_contact; rows matching an existing contact on email, UID, or name+postcode come back in duplicates[] (never auto-merged) so the operator resolves each. Returns {created, createdIds, skipped, duplicates}.
| Name | Required | Description | Default |
|---|---|---|---|
| rows | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral disclosure burden. It does well by specifying duplicate matching fields (email, UID, name+postcode), the 'never auto-merged' rule, and the exact return shape. However, it omits essential behaviors around the required idempotencyKey semantics, what causes a row to be 'skipped', and whether the import is atomic or partial on failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core action ('Import already-parsed contact rows with duplicate detection') and then efficiently adds the most important behavioral details. No fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without an output schema or annotations, the description must fully cover return semantics and edge cases. It lists the return keys (created, createdIds, skipped, duplicates) but doesn't define what 'skipped' means, how duplicates are represented (IDs, objects?), or how idempotencyKey affects the operation. This leaves critical operational gaps for an import tool that can create or skip many records.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the description only hints at row content through the duplicate criteria (email, UID, name+postcode). It does not explain the structure of `rows`, the purpose of `workspaceId`, or the role of `idempotencyKey`. The reference to 'validated like create_contact' is indirect and doesn't compensate for the complete lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Import') and resource ('already-parsed contact rows'), and distinguishes itself from single-record create_contact by explicitly mentioning duplicate detection and the 'never auto-merged' policy. This makes it clear what the tool does and how it differs from nearby siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than explicit: the description implies bulk import of pre-parsed rows polls, but it never names alternatives like create_contact or contacts_merge, nor does it state when to prefer this tool or list exclusions such as 'do not use for a single contact'. The context is present, but the guidance is not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_log_activityA
Log a note, call, email, meeting or task on a contact (the OP5 activity seam). occurredAt may be backdated but never future-dated; the stream is append-only (a wrong note is corrected by a new note). kind is one of note|call|email|meeting|task.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | ||
| kind | Yes | ||
| dealId | No | ||
| contactId | Yes | ||
| occurredAt | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and delivers: occurredAt may be backdated but never future-dated, the stream is append-only with corrections made by new notes, and kind is restricted to the listed five values. It does not cover idempotencyKey retry semantics or permission requirements, but the most decision-relevant behavioral traits are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with zero filler: core action first, then temporal constraint, then kind enumeration. Each sentence carries distinct, load-bearing information and the structure front-loads the purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter write tool with no annotations and no output schema, the description covers the critical semantics (append-only, time constraints, kind values) but omits the idempotencyKey behavior that matters for retry-safe writes, gives no guidance on optional dealId usage, and doesn't state what a successful call returns. The essential calling context is present, but meaningful gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it adds meaning for two key parameters: kind (enumeration note|call|email|meeting|task, absent from the schema's plain string) and occurredAt (backdatable, not future-dated). However, idempotencyKey's dedupe role, dealId's optionality, and body's content expectations are left unexplained, leaving several parameters underspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ("Log a note, call, email, meeting or task on a contact") with a clear verb and resource, and enumerates the activity kinds. The parenthetical "(the OP5 activity seam)" is unexplained jargon that slightly obscures, and it doesn't explicitly contrast with sibling tools like contacts_timeline, but the core purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied by the verb and resource — an agent can infer this is the write-seam for contact activity. However, no alternatives or exclusions are named (e.g., contacts_timeline for reading activity, deals_log_activity for deal-scoped activity), and there's no explicit statement of when to choose this tool over them. The append-only note hints at boundaries but doesn't route between siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_mergeA
Merge a duplicate contact into a survivor: re-points every live foreign key that names the source (documents, activities, employer links) to the target and tombstones the source. No posting and no journal effect. It does NOT freeze what an already-issued invoice re-renders as: A11 keeps no snapshot and reads the live contact row, so the survivor identity is what a re-rendered QR-bill carries. When the two roles differ the survivor is promoted to partyRole "both" and the result says so. Idempotent; a self-merge or a tombstone target is refused.
| Name | Required | Description | Default |
|---|---|---|---|
| sourceId | Yes | ||
| targetId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so thoroughly: it discloses the re-pointing of every live foreign key, tombstoning of the source, no posting/journal effect, idempotency, refusal conditions, and the survivor promotion to partyRole 'both' when roles differ. This goes far beyond what the schema or annotations could convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: scope, side effects, exclusions, role promotion, and refusal conditions. It is front-loaded with the core action and then layers caveats. Slightly long, but the complexity of the merge operation justifies the length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers the essential behavioral context: what gets re-pointed, what gets tombstoned, what does NOT happen (no posting/journal effect), idempotency, refusal cases, and role promotion. An agent has enough to call it correctly and to anticipate the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the semantic roles of sourceId and targetId ('duplicate contact' vs 'survivor') and the workspaceId context, and mentions idempotencyKey implicitly via 'Idempotent'. It doesn't spell out the exact format of each parameter, but the merge semantics are clear enough for an agent to map sourceId/targetId correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Merge a duplicate contact into a survivor') and immediately distinguishes the operation from a simple update by stating the re-pointing of foreign keys and tombstoning of the source. It also names the sibling it is not (contacts_merge vs contacts_anonymise, contacts_archive, etc.) by describing the exact merge semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool (merging duplicates) and what it does NOT do ('does NOT freeze what an already-issued invoice re-renders as'), which is a clear exclusion. It also names the alternative behavior (A11 keeps no snapshot and reads the live contact row) and states refusal conditions (self-merge or tombstone target), giving an agent enough to decide when to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_tagA
Tag a contact: merge role and segment values into its roles/segments (deduplicated, case-preserving). Unlike update_contact, which sets the arrays absolutely, this UNIONS the new values into the existing set. Tags are free-text (trimmed, at most 40 chars each, at most 50 per contact), never a fixed enum.
| Name | Required | Description | Default |
|---|---|---|---|
| roles | No | ||
| segments | No | ||
| contactId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure, and it succeeds: it states the operation mutates by merging/unioning, preserves case, deduplicates, trims free-text tags, enforces per-tag length limits, per-contact limits, and clarifies tags are not enum-based. These are concrete behavioral expectations beyond the raw schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler: the first states the action, the second delivers the critical sibling distinction, and the third covers constraints. The information is front-loaded and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating tool with no annotations, no output schema, and 0% schema description coverage, the description provides strong operational context: main semantics, the distinction from update_contact, and validation constraints. Minor gaps remain around return behavior and idempotencyKey semantics, but nothing prevents an agent from invoking the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate; it does for the key parameters by explaining roles and segments are arrays of free-text tags with trimming, case preservation, and size limits. It does not explain idempotencyKey or explicitly describe workspaceId/contactId, though their roles are largely evident from their names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Tag a contact' with a precise behavior: merging role and segment values into existing roles/segments. It also explicitly differentiates itself from update_contact, which sets the arrays absolutely rather than unioning. This makes the tool's purpose unmistakable and distinguishable from its siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly names update_contact as the alternative and contrasts the two behaviors: contacts_tag UNIONs values while update_contact sets arrays absolutely. This gives an agent clear conditions for when to choose this tool versus a sibling. No further alternative routing is needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contacts_timelineARead-only
Read a contact's activity timeline, newest first (P5). Following a merge chain, a read of a merged-away contact returns the survivor's consolidated timeline.
| Name | Required | Description | Default |
|---|---|---|---|
| contactId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the 'newest first' ordering and the merge-chain behavior, which are useful behavioral traits beyond the readOnlyHint annotation. It does not contradict the annotation and adds context about what happens with merged contacts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no fluff. It front-loads the primary purpose and then adds a relevant edge-case behavior. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with two parameters and no output schema, the description covers ordering and merge behavior. It does not mention pagination or limits, but these are minor for a read operation. Overall, it provides sufficient context for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero description coverage for both parameters. The description mentions 'a contact's timeline', implicitly referencing contactId, but never explains workspaceId or any parameter format or constraints. It fails to compensate for the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb and resource: 'Read a contact's activity timeline, newest first'. It also adds behavioral detail about merge chains. However, it does not explicitly differentiate from sibling tools like contacts_tag or contacts_log_activity, though the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for reading a contact's timeline but provides no explicit guidance on when to choose this over alternatives (e.g., contacts_log_activity). The merge-chain note hints at a specific scenario but does not serve as a clear usage guideline or exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
convert_documentC
Convert an accepted quote to an order or invoice (or a confirmed order to an invoice), linking the source.
| Name | Required | Description | Default |
|---|---|---|---|
| toType | Yes | ||
| documentId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It implies a mutation (converting documents) but does not disclose side effects, such as whether the original document is modified or superseded, permissions required, reversibility, or any consequences for linked records. The phrase 'linking the source' hints at association but lacks detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that states the core functionality without unnecessary words. It is well-structured and front-loaded with the primary action. However, its brevity contributes to the lack of behavioral and parameter details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with four parameters, no annotations, and no output schema, the description is severely incomplete. It does not specify prerequisites (e.g., quote must be 'accepted', order must be 'confirmed'), valid toType values, the meaning of 'linking the source', or any post-conversion effects. This leaves significant gaps for an agent attempting to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain any of the four parameters (toType, documentId, workspaceId, idempotencyKey). Notably, toType is a string with no enum, so allowed values (e.g., 'order', 'invoice') are undocumented. The description adds no semantic value beyond what the parameter names imply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: converting an accepted quote to an order or invoice, or a confirmed order to an invoice, with linking. It names specific source and target document types, which gives a clear verb+resource. However, it does not explicitly differentiate from sibling tools like quotes_convert or sales_order_from_quote, leaving some ambiguity about when this tool is the right choice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives. Sibling tools such as quotes_convert, sales_order_from_quote, and sales_order_invoice likely overlap in functionality, but the description does not mention them or provide conditions for selection. Users are left to infer usage from the name and description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
costing_budget_vs_actualARead-only
Budgetvergleich (budget vs actual) for one project, in the workspace base currency: the B00 budget (the H-FX base snapshot where one was taken) against B03's cost-to-date (time at snapshot rates plus posted project bills plus the received-not-billed purchase accrual, the same terms B00's own budget seam reports), with remainingMinor, consumedBp (basis points, rounded once), logged hours vs budgetHours, and the overBudget flag. A project without budget fields answers budgeted:false and cost-to-date only, never a fake 0-budget overrun. remainingMinor goes negative on overrun, which an automation rule can read as condition data to fire an existing action (B03 emits no events of its own).
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| projectId | Yes | ||
| workspaceId | Yes | ||
| includeOpenTime | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description reveals important behavior: no fake zero-budget overrun, negative remainingMinor on overrun, rounding of consumedBp, inclusion terms for cost-to-date, and the absence of B03 events. This gives an agent reliable expectations for what a successful read will and will not do.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and packs many useful details without filler, and the core purpose appears first. It is long and somewhat parenthetical, but every sentence adds needed domain context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by listing the returned fields and key edge cases, and it explains currency and automation integration. It is complete for the required call path, though the optional parameter semantics and lack of an explicit return shape keep it just short of fully comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but the description compensates by explaining workspaceId's base-currency context and projectId's project-level scope, while also hinting at temporal aspects through snapshot and cost-to-date language. However, asOf and includeOpenTime are never explicitly explained, so those two optional parameters still rely on name inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that this returns a budget-vs-actual comparison for a single project in the workspace base currency, and it names the specific output metrics such as remainingMinor, consumedBp, and overBudget. It is precise about scope and calculation, though it does not explicitly distinguish itself from close siblings like project_budget_actual.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It communicates a concrete automation use case (reading remainingMinor as condition data) and explains behavior on missing budget fields, so an agent can infer when the tool is relevant. However, it never names alternatives or gives explicit when-to-use vs. when-not-to-use guidance relative to the many costing and reporting siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
costing_drilldownARead-only
Die Belege hinter dem Projekterfolg: the raw contributing rows for ONE component ('time' | 'expenses' | 'purchases' | 'accrued_purchases' | 'committed' | 'revenue'), so every Rappen of the margin resolves to its source: time entries with snapshot rates and minutes, posted vendor bills with their postedEntryId, PO lines with quantity and base unit price, posted invoice lines and credit-note lines with their postedEntryId (the OR 957a traceability chain into the journal). Keyset-paginated (cursor/nextCursor, limit up to 500); totalMinor is the whole component and equals the card figure exactly, same code path. A component whose source rows carry no project reference answers zero rows with unattributable:true (none does today). An unknown component answers invalid_component.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| basis | No | ||
| limit | No | ||
| cursor | No | ||
| groupBy | No | ||
| component | Yes | ||
| projectId | Yes | ||
| workspaceId | Yes | ||
| includeOpenTime | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only readOnlyHint=true in annotations, the description adds substantial behavioral context: keyset pagination via cursor/nextCursor, limit up to 500, totalMinor matching the card figure exactly, same code path guarantee, the unattributable:true edge case, and invalid_component for unknown components. No contradiction with the readOnly annotation exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and almost every clause adds a distinct fact, including pagination, traceability, and edge-case behavior. It is slightly run-on and mixes languages, and contains an odd 'OR 957a' fragment, but it remains well front-loaded and free of fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only drilldown, the description covers the main behavioral surface: source-row kinds, pagination, totals, and error/edge cases. Without an output schema it could go further on exact response field names or the meaning of ambiguous parameters like `basis` and `groupBy`, but the essential invocation context is largely present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does document the central `component` parameter's allowed values and explains `cursor`, `nextCursor`, and `limit`. However, it leaves several parameters meaningfully unexplained, including `asOf`, `basis`, `groupBy`, and `includeOpenTime`, and never clarifies the required `workspaceId`/`projectId` semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific operation: returning raw contributing rows for one of six enumerated components, and gives concrete examples of source rows (time entries, vendor bills, PO lines, invoice lines). It is clearly distinguishable from the costing_project_pl/costing_budget_vs_actual siblings by focusing on row-level drilldown rather than period totals, though it never explicitly names those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description conveys when this tool is useful: to make every margin figure resolvable to source documents and to get a component-specific drilldown. However, it provides no explicit guidance about when to use an alternative like costing_project_pl or costing_pl_list, and no 'do not use when' conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
costing_pl_listARead-only
Projekterfolg-Portfolio (P&L list): one revenue/cost/margin summary row per project, margin-sorted (die Marge, descending), the read that answers "welche Projekte oder Mandate tragen?". Closed projects are excluded unless status='closed' asks for them; any single B00 status filters exactly. A project with foreign-currency rows degrades to fxBaseMissing:true instead of blanking the portfolio. savedViewId applies a G00 saved view over the project entity kind: its stored filters merge underneath any filter named explicitly here. Same computation and same honesty flags as costing_project_pl (basisDegraded when a cost-basis entry lacks a cost rate).
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| basis | No | ||
| status | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| includeOpenTime | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses significant non-obvious behaviors: margin-descending sort, closed-project exclusion unless status='closed', exact B00 filtering, degradation to fxBaseMissing for foreign-currency rows, saved view filter merging, and the basisDegraded honesty flag shared with costing_project_pl. These details substantially exceed the annotation baseline and materially aid an agent in interpreting results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-rich, covering many edge cases, but it mixes German and English and uses cryptic codes (B00, G00) without explanation. The structure is logical—output shape, filtering, degradation, saved views, shared computation—but the phrasing is not lean and some sentences are overloaded. It earns its place but sacrifices readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (6 parameters, no output schema, no parameter descriptions), the description covers the output shape and several behavioral edge cases, but leaves core parameters undefined and does not describe the exact return fields. An agent can call it correctly for the covered filters but may misuse asOf, basis, or includeOpenTime. Incomplete but not severely so.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden for all 6 parameters. It explains only status (via closed/B00 semantics) and savedViewId (merge behavior). The parameters asOf, basis, includeOpenTime, and workspaceId are never mentioned, leaving an agent guessing about their meaning and accepted values. This is a major gap for a tool with no schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it reads a portfolio of project P&L summaries, one row per project, sorted by margin descending. It directly answers the business question 'welche Projekte oder Mandate tragen?', and references sibling costing_project_pl to place itself as the portfolio-level version of the same computation, distinguishing it without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context: this is the read for portfolio-level project profitability, with explicit filter semantics for closed projects, B00 status, and saved view merging. It does not explicitly name alternatives or say when not to use it, but the referenced costing_project_pl and the portfolio framing imply the boundary. No contrary usage guidance is present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
costing_project_plARead-only
Projekterfolg (project P&L card): revenue vs cost for ONE project in integer Rappen, recomputed live from the source rows on every call (pure read, B03 posts nothing and caches nothing). Cost splits into time | expenses | purchases | accrued_purchases: 'time' is B01 entries at their OP1 snapshot rate (approved and beyond; includeOpenTime widens to open/submitted WIP), 'expenses' and 'purchases' are POSTED project-tagged A17 vendor bills at stored base net (purchases when 3-way matched to a D02 PO, expenses otherwise), 'accrued_purchases' is received-not-yet-billed project-tagged D02 quantity at PO base price (it drops as the matched bill posts, so no Rappen counts twice). committedMinor (ordered minus received on open POs) reports BESIDE the cost, never inside it. Revenue counts POSTED invoice lines generated from this project's time (B02's linkage), net of posted credit notes; drafts never count. marginBp is null when revenue is 0, never a fake break-even. basis 'bill' or 'cost': 'cost' values time at the rate card's cost-rate snapshot and falls back per entry to the bill rate with basisDegraded true where none was defined. asOf cuts every component by its source date. groupBy re-buckets the time component by a confirmed select custom field on time_entry without moving any total. A row priced in a non-base currency answers fx_base_missing rather than mixing currencies.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| basis | No | ||
| groupBy | No | ||
| projectId | Yes | ||
| workspaceId | Yes | ||
| includeOpenTime | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true, so the read-only behavior is known. The description adds significant context about what gets computed and excluded, such as 'pure read, B03 posts nothing and caches nothing', which goes beyond annotations. However, it doesn't clarify auth or rate limits, but given the readOnlyHint, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-packed, but it's long. It front-loads the core purpose and read-only nature, with detailed explanations of cost components and parameters. Each sentence adds value, so it's appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and lack of output schema, the description covers the key semantics: cost splits, revenue handling, margin behavior, basis options, and parameter effects. The main gap is lack of return structure, but without an output schema, the description compensates well.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description explains several parameters in depth: 'basis' (bill vs cost), 'includeOpenTime' (widens to open/submitted WIP), 'asOf' (cuts components by source date), 'groupBy' (re-buckets time). It does not explicitly explain 'workspaceId' and 'projectId', but those are obvious. This compensates for the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the tool computes the project P&L (revenue vs cost) for ONE project, details the cost splits, specifies that it's a pure read, and differentiates it from siblings like costing_drilldown and costing_budget_vs_actual by mentioning 'ONE project' and 'recomputed live from source rows'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the behavior in detail but does not explicitly state when to use this tool vs alternatives like costing_drilldown or costing_pl_list. It implies usage (for a single project P&L card) but doesn't name alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_accountC
Create a ledger account.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| type | Yes | ||
| number | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No | ||
| vatCodeDefault | No | ||
| costCenterAllowed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only states that an account is created, with no information about validation, idempotency, duplicate-account handling, required permissions, side effects, or response behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The sentence is short, front-loaded, and free of filler, which is good for structure. However, for a 7-parameter creation tool with no schema descriptions and no annotations, the description is under-sized and does not provide enough substance to be considered appropriately concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 7 parameters, no parameter descriptions, no output schema, no annotations, and no usage context, the description is severely incomplete. An agent has almost no information about required account attributes, valid account types, idempotency semantics, or what a successful response looks like.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for 7 parameters, and the description adds no meaning beyond the word 'ledger account'. It does not explain workspaceId, number, name, type, idempotencyKey, vatCodeDefault, or costCenterAllowed, leaving the agent without enough information to populate them correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Create') and resource ('a ledger account'), which clearly differentiates it from sibling tools like create_bank_account, update_account, archive_account, and list_accounts. An agent can immediately understand the tool's core function from this one line.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives such as update_account or create_bank_account. It also does not mention prerequisites like an existing workspace, chart-of-accounts context, or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_automation_ruleA
Define an automation: when this event happens, and this condition holds, call this write verb with this input. The action must name a registered write verb, the trigger must name a registered event, and a rule whose action is the very verb that emits its own trigger is refused outright. A rule created by the agent actor lands DISABLED and a human must enable it.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| action | Yes | ||
| enabled | No | ||
| trigger | Yes | ||
| condition | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses important behaviors: registration validation, outright refusal of self-referential rules, and the disabled-by-default policy for agent-created rules. It omits permissions, idempotency, and return behavior, but the most consequential side effect is clearly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences contain no filler. The first sentence front-loads the core definition, the second imposes validation constraints, and the third states the actor-specific default. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter tool with nested objects and no output schema, the description provides a solid conceptual map and important caveats. Yet it leaves the exact JSON shapes of action and trigger unspecified and does not mention where to discover registered verbs/events or explain idempotency/workspace fields. It is minimally adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates partially by explaining the roles of trigger, condition, and action, and by noting the disabled default. It does not clarify workspaceId, name, enabled, or idempotencyKey beyond that, leaving some parameter semantics unspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool defines an automation rule by combining trigger, condition, and action. It conveys creation semantics and key structural elements, though it does not explicitly contrast with sibling tools like update_automation_rule or enable_automation_rule.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains prerequisites and constraints: the action must name a registered write verb, the trigger must name a registered event, self-referential rules are refused, and agent-created rules are disabled by default. However, it never points to alternatives (e.g., update or enable tools), so the when-to-use guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_backupA
Take a point-in-time .tillbackup snapshot: a tenant-scoped SQLite file plus a re-hashed manifest, recorded in the backup history. The byte-perfect artifact restore_backup consumes. An optional planId links the backup to a Datenübernahme plan (once it has reached planned), satisfying the migration commit gate’s pre-migration-backup leg; a planId in another workspace refuses.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does substantial work: it describes the produced artifact, the recorded history, byte-perfect restorability, plan-state requirement, and cross-workspace refusal. It stops short of explaining idempotencyKey behavior (e.g., duplicate prevention on retry) or authorization needs, which are the main gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences front-load the core action and artifact, then add the plan-linkage constraint. Every sentence contributes distinct, non-redundant information and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating backup tool with no annotations and no output schema, the context is strong but incomplete: it covers the artifact and plan-linkage rule but omits idempotencyKey semantics, return/error behavior, and permissions. An agent could call it, but a key required parameter remains under-specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It richly explains planId: optional, links to a Datenübernahme plan only after planned, and refuses when in another workspace, and 'tenant-scoped' implies workspaceId's role. However, the required idempotencyKey parameter is never explained, leaving a significant semantic gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and resource: 'Take a point-in-time .tillbackup snapshot' of a tenant-scoped SQLite file plus re-hashed manifest. It further distinguishes the tool by identifying the artifact as what restore_backup consumes, separating it from backup listing/deletion/verification siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives concrete context: backups are recorded in backup history, are restorable, and planId should be supplied when the migration commit gate requires a pre-migration backup and the plan has reached planned. It also warns that a planId in another workspace is refused, but it does not explicitly contrast create_backup with alternatives such as export_workspace or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_bank_accountA
Register a Bankkonto: validates the IBAN (ISO 13616 mod-97), derives whether it is a QR-IBAN (QR-IID 30000 to 31999, receive-only), and links it to an asset account in the chart (ledgerAccountId, e.g. 1020 Bankkonto). Does NOT set an opening balance: that is set_bank_opening_balance.
| Name | Required | Description | Default |
|---|---|---|---|
| iban | Yes | ||
| name | Yes | ||
| currency | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No | ||
| ledgerAccountId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It clearly states validation (IBAN mod-97), QR-IBAN derivation with specific IID range and receive-only semantics, and the linking action. It also explicitly calls out a non-behavior (no opening balance), which is valuable. It omits error handling, permissions, and return format, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The main action is front-loaded, and the exclusion is placed at the end for emphasis. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the tool's main behavior and one critical exclusion, but it lacks parameter semantics for most parameters and provides no output schema or return-value information. For a create operation with no annotations, this is incomplete, though the core purpose is clear enough to call correctly for the primary use case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains 'iban' (validation) and 'ledgerAccountId' (with example '1020 Bankkonto'), but leaves 'workspaceId', 'name', 'currency', and 'idempotencyKey' completely unexplained. This is a significant gap for a 6-parameter tool, especially since four are required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource, 'Register a Bankkonto', and explains the core behavior: IBAN validation, QR-IBAN detection, and linking to a ledger account. It also distinguishes itself from the sibling set_bank_opening_balance by explicitly stating it does not set an opening balance, making the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states what this tool does NOT do ('Does NOT set an opening balance') and names the alternative tool for that action ('that is set_bank_opening_balance'). This gives clear when-to-use and when-not-to-use guidance, though it doesn't discuss update_bank_account, the exclusion is highly relevant.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_contactC
Create a customer/vendor contact.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| No | |||
| address | No | ||
| partyRole | Yes | ||
| vatNumber | No | ||
| description | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No | ||
| defaultCurrency | No | ||
| paymentTermsDays | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden. 'Create' implies a mutating/persisting operation, but the description does not mention idempotency, duplicate handling, validation rules, side effects, or return behavior, leaving an agent without enough transparency for a create operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler or redundancy. It achieves maximum conciseness, even though this brevity sacrifices detail that is penalized in other dimensions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter create tool with a nested address object and no output schema, one sentence is severely incomplete. It omits required inputs, object structure, idempotency behavior, and expected results, making it insufficient for an agent to invoke correctly with confidence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not compensate. It offers only a vague hint about partyRole via 'customer/vendor' and says nothing about required fields like workspaceId, the nested address object, name, vatNumber, idempotencyKey, or paymentTermsDays.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Create'), a resource ('contact'), and the subtype scope ('customer/vendor'). It is clear and distinguishable from sibling read/update/archive tools by the create action, though it does not explicitly name any alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus related tools such as update_contact, contacts_merge, or contacts_import. No preconditions, exclusions, or alternative selection criteria are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_cost_centerC
Create a cost centre (Kostenstelle).
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | ||
| name | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden for behavioral disclosure, and it only restates the creation action. It does not mention uniqueness constraints on code, conflict behavior on duplicates, required permissions, side effects, or retry/idempotency semantics despite an idempotencyKey parameter being present. The presence of idempotencyKey hints at retry semantics that the description never explains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with zero filler words, and the domain translation is front-loaded. The conciseness, however, reflects under-specification rather than efficient packing of essential decision-making information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a creation tool with 4 undocumented required parameters, no annotations, and no output schema, the one-sentence description is severely inadequate. An agent cannot determine what a valid code looks like, which workspace context is needed, or what to expect in the response. Nothing beyond the basic action is conveyed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description names none of the four parameters (workspaceId, code, name, idempotencyKey). An agent cannot distinguish what 'code' means versus 'name' or how idempotencyKey should be used, since neither the schema nor the description provides definitions. This is the largest gap: the tool adds absolutely no semantic meaning beyond the raw parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Create a cost centre'), so the core action is unambiguous. However, it is a near-restatement of the name 'create_cost_center with only the German translation (Kostenstelle) adding genuine semantic value. It neither explicitly differentiates from siblings, though the action is clear enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool over alternatives such as list_cost_centers, archive_cost_center, or delete_cost_center. No preconditions, sequencing, or contextual triggers for when creation is appropriate are mentioned. An agent must infer usage entirely from the name with no help.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_credit_noteA
Draft a Gutschrift against an issued invoice: full, selected positions (reduced quantities), or a net amount apportioned over the remaining line nets. Always a draft; issuing is a separate, confirmed step (P8).
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | ||
| lines | No | ||
| reason | No | ||
| amountMinor | No | ||
| workspaceId | Yes | ||
| fromInvoiceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the core behavior (always a draft, not issuing) and that it operates against an issued invoice. However, it does not mention side effects, reversibility, idempotency requirements, or any failure behavior—gaps that would be valuable for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the main purpose and then the mode options and draft behavior. No fluff or redundancy—every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 7 parameters, no output schema, and no annotations, the description covers the core usage (modes and draft nature) but omits details like the idempotencyKey requirement and why it's needed, plus the reason parameter. It is adequate for basic use but not fully complete for an agent that must handle all required fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the three modes (full, selected positions, net amount), which maps to mode, lines, and amountMinor. However, it does not explain idempotencyKey, reason, workspaceId, or fromInvoiceId, leaving several required parameters underspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Draft), resource (Gutschrift/credit note), and target (against an issued invoice), and clearly distinguishes the three modes (full, selected positions, net amount apportioned). It also separates drafting from issuing, which is a key distinction from the sibling issue_credit_note tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says 'Always a draft; issuing is a separate, confirmed step (P8)', which tells the agent when to use this tool (to draft) and when not to (to issue). It doesn't name the exact sibling issue_credit_note, but the guidance is clear enough to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_demo_workspaceA
Mint a disposable demo workspace (workspace.kind = demo) and seed it with sample Swiss books through the real verbs: chart auto-seed, MWST effektiv/Soll, three customers, three items, two issued invoices and one draft, every posting through issue_invoice, so the demo behaves identically to a real workspace. The caller is seated as owner. No MWST number or UID is fabricated; the creditor profile carries the SIX specimen QR-IBAN. A demo can never go productive: its only exit is discard_demo_workspace. Idempotent: retrying the same key returns the existing demo, never a second one.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden, and it delivers: it discloses the seeding process through real verbs, the caller's owner role, the absence of fabricated MWST/UID data, the SIX specimen QR-IBAN, the inability to go productive, and idempotency. This gives an agent a strong model of side effects and constraints without contradicting any structured metadata.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause earns its place—setup scope, data details, key constraints, exit path, and idempotency. The most important usage signal is front-loaded, and the exclusion of production use appears early.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description covers an impressive amount: what gets seeded, how, ownership, creditor details, lifecycle constraints, and retry semantics. The only meaningful gap is that it never states the return value or output shape, and the 'name' parameter is left unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clearly explains idempotencyKey through the phrase 'retrying the same key returns the existing demo.' However, the 'name' parameter is entirely unaddressed in both the schema and the description, leaving its purpose and optionality to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair ('Mint a disposable demo workspace') and immediately differentiates it from real workspaces by noting workspace.kind = demo and the fact that it can never go productive. It also names the exit path (discard_demo_workspace), which distinguishes it from sibling tools like create_workspace and bootstrap_workspace.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes clear this is for disposable demos by stating the workspace 'can never go productive' and that its only exit is discard_demo_workspace. It also gives the idempotent retry behavior. It does not explicitly name alternatives such as create_workspace or bootstrap_workspace for real/setup scenarios, so it stops short of a full when-to-use/when-not-to-use guide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_documentC
Create a document draft (quote, order, invoice, or credit note) with positions.
| Name | Required | Description | Default |
|---|---|---|---|
| type | Yes | ||
| lines | No | ||
| notes | No | ||
| dueDate | No | ||
| currency | No | ||
| contactId | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It only says a draft is created with positions, but does not disclose whether the draft is persisted immediately, whether it can be issued later, what validation occurs, what the response contains, or whether there are side effects. 'Draft' gives a hint but not enough context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single clear sentence with no filler. It front-loads the core action and resource, then adds key scope details. Every word contributes value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters, no output schema, and no annotations, this description is not sufficient for an agent to invoke the tool correctly. It omits required parameter semantics, line item requirements, drafting behavior, and return information. The description is a useful summary but not operationally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does hint that 'positions' maps to the lines parameter and that type can be one of quote/order/invoice/credit note. However, it provides no meaning for workspaceId, contactId, dueDate, currency, notes, or idempotencyKey, leaving most parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Create a document draft') and defines the resource scope with explicit document types: quote, order, invoice, or credit note. The word 'draft' also distinguishes this from finalizing or issuing tools, though it does not directly contrast with siblings like create_credit_note or issue_invoice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternative document-creation tools such as quotes_create, sales_order_create, create_credit_note, or issue_invoice. The 'draft' wording implies a use case, but no explicit when/when-not or alternative routing is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_document_templateA
Create a document template (invoice, credit_note, quote or dunning_run): footer text per locale, render language (fixed or per contact), an ordered set of OPTIONAL non-legal line-item columns, and an E00 logo file linked via entityKind document_template. Always saved as a NON-default draft: going live is a separate set_default_document_template call a human confirms. Presentation only: the Swiss QR-bill payload, VAT figures and legal content are produced by the document itself and pass through byte-identical under every template.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| footerI18n | No | ||
| fixedLocale | No | ||
| workspaceId | Yes | ||
| documentKind | Yes | ||
| languageMode | No | ||
| idempotencyKey | No | ||
| logoDocumentId | No | ||
| lineItemColumns | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that the template is always saved as a non-default draft, that legal content and QR-bill payload are unaffected, and that the logo is linked via entityKind document_template. This provides meaningful behavioral context beyond the raw schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and well-structured. It front-loads the core purpose, then clarifies the draft behavior and the presentation-only nature. Each sentence adds value without redundancy, though it is somewhat long.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (9 params, nested objects, no output schema), the description provides a good high-level overview but misses critical details: it does not explicitly list required parameters, does not describe idempotencyKey behavior, and does not indicate what the response looks like. It does clarify the draft/live distinction and the document kinds, which is helpful, but for a tool with zero schema coverage and no output schema, more parameter-level detail would be needed for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameters. It does cover several: footerI18n (footer text per locale), languageMode/fixedLocale (fixed or per contact), lineItemColumns (ordered optional non-legal columns), logoDocumentId (E00 logo file), and documentKind (enumerated in parentheses). However, it omits semantics for required params like workspaceId and name, and for idempotencyKey. The description adds value but does not fully compensate for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states what the tool does: creates a document template for specific document kinds (invoice, credit_note, quote, dunning_run) with configurable footer, language mode, optional line-item columns, and logo. It distinguishes from sibling set_default_document_template by explicitly stating it always creates a draft, not a live template.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear guidance on when to use this tool vs. alternatives: it explicitly says going live is a separate set_default_document_template call and that the template is presentation-only, implying it should not be used for legal content changes. It does not mention update_document_template for editing existing templates, but the draft/live distinction is clearly articulated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_entry_for_txnA
Book an unmatched txn (a bank fee, interest, a standing transfer) as a balanced two-leg journal entry via A02 (the bank leg plus the chosen contra account) and link it. Refuses currency_mismatch for a txn whose currency differs from the workspace base, and refuses taxCode (declared for a future increment, not modelled yet): book the net and add VAT as a separate manual entry. CONSEQUENCE: Books an unmatched bank transaction as a journal entry; the only correction is a reversing entry.
| Name | Required | Description | Default |
|---|---|---|---|
| taxCode | No | ||
| bankTxnId | Yes | ||
| description | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| contraAccountId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full disclosure burden and succeeds: it reveals the two refusal modes (currency mismatch vs workspace base; unsupported taxCode), the workaround for taxCode, and the critical consequence that the only correction is a reversing entry. This is exceptional transparency for a mutation tool with no annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded: the main purpose and mechanism land in sentence one, refusals and workaround in sentence two, and consequence in sentence three. Nothing is wasted, though the CONSEQUENCE line partially restates the purpose ('Books an unmatched bank transaction as a journal entry') and the taxCode parenthetical is slightly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a moderate-complexity mutation with no annotations and no output schema, the description covers purpose, mechanism, failure modes, workaround, and correction path. The only notable gaps are the response/success signal (no output schema exists) and explicit idempotency behavior, but the operation is specified well enough for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does for the key parameters: contraAccountId is explained as 'the chosen contra account', bankTxnId maps to the unmatched txn being booked, and taxCode semantics are fully detailed (declared but not modelled, must be refused). workspaceId, idempotencyKey, and description are left to naming conventions, which is acceptable but keeps this from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Book'), resource ('unmatched txn'), and mechanism ('balanced two-leg journal entry via A02: bank leg plus chosen contra account'), then states it links the entry to the transaction. It is clearly distinguishable from siblings like post_entry (general journal entry) and reverse_entry (correction) without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool — unmatched bank transactions such as bank fees, interest, and standing transfers — and states exclusions: it refuses currency_mismatch and refuses taxCode, directing the agent to book the net amount and add VAT via a separate manual entry. It stops short of naming the alternative tool explicitly, but the workaround is actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_itemC
Create an invoicing item (a reusable line default).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| unit | No | ||
| currency | No | ||
| workspaceId | Yes | ||
| defaultTaxCode | No | ||
| idempotencyKey | No | ||
| revenueAccountId | No | ||
| defaultUnitPriceMinor | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses that the operation creates a resource and that the item is reusable, but it does not explain idempotency semantics, side effects, permission needs, or what happens on success. This is minimal disclosure for a creation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence with no filler; the action is front-loaded and the parenthetical adds clarifying value. It is concise, though the brevity verges on under-specification for a tool with 8 parameters and no annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, no annotations, and 8 undocumented parameters, yet the description only states the core purpose. It does not explain required fields, idempotencyKey behavior, return values, or prerequisites. An agent cannot reliably construct a correct call from this description alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description names none of the 8 parameters. The parenthetical 'reusable line default' adds some meaning to defaultUnitPriceMinor/defaultTaxCode, but workspaceId, name, unit, currency, revenueAccountId, and idempotencyKey remain unexplained. The description does not compensate for the schema's lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Create') and resource ('invoicing item'), and the parenthetical '(a reusable line default)' clarifies what kind of entity this is. This clearly distinguishes it from generic create tools and from update_item/archive_item lifecycle siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or alternative guidance is provided. The only usage signal is the imperative 'Create,' so an agent must infer that this should be used for new item creation and that update_item/archive_item handle other lifecycle stages. No prerequisites, exclusions, or context are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_payment_batchA
Draft a payment batch from open A17 bills: validates the debtor account (A19, must not be a QR-IBAN, which is receive-only), that every bill is posted and open, that every bill shares one currency (CHF or EUR; a mixed selection is refused with mixed_currency), that every vendor has a creditor_bank_profile, and that a QR-IBAN vendor has a valid QRR reference on the bill. Snapshots each item so a later generate_pain001 is byte-reproducible even if the vendor profile changes afterward. Posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| itemIds | Yes | ||
| workspaceId | Yes | ||
| bankAccountId | Yes | ||
| executionDate | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden and does well: it discloses the validation steps, the snapshot behavior for reproducibility, the refusal of mixed currency with a specific error code, and explicitly states it posts nothing. This is comprehensive behavioral disclosure for a non-mutating draft action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph but each sentence adds value. It front-loads the main purpose and then details validations and snapshot behavior. It is fairly long but not redundant; no fluff. Could be slightly more structured with bullet points, but effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (multiple validations) and lack of output schema, the description covers the essential behaviors and the main failure case (mixed_currency). It does not mention other error codes or the exact return value, but for a draft operation that posts nothing, the provided details are sufficient for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it does not explain any of the five parameters. It never mentions what itemIds are (presumably the bills), the purpose of executionDate, or idempotencyKey semantics. The tool description focuses on validation logic, leaving parameter meaning entirely to the agent's inference from names, which is insufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the verb (draft), resource (payment batch), and source (open A17 bills). It also distinguishes itself from later stages like generate_pain001 by noting it posts nothing. The scope is precise and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: it is the drafting step before generate_pain001, and it validates specific conditions. However, it does not explicitly name alternatives or say when NOT to use it (e.g., if bills are not open). The mention of 'a later generate_pain001' implies the workflow but doesn't contrast with other payment-related siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_recurring_scheduleA
Lege eine Serienrechnung an (a recurring invoice schedule): the template (contact + positions, or templateDocumentId to snapshot an existing document), the cadence (interval monthly|quarterly|yearly|custom with customDays, anchorDate ISO YYYY-MM-DD), and optionally endDate, maxOccurrences, dueDays payment terms and autoIssue. Nothing is invoiced yet: run_due_recurring materialises each due period, as a review draft unless autoIssue is on (P8). The template is SNAPSHOTTED, never a live link, and never stores a supply date: the tick stamps each generated line with its own period as the Leistungsdatum.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| lines | No | ||
| notes | No | ||
| dueDays | No | ||
| endDate | No | ||
| currency | No | ||
| interval | Yes | ||
| autoIssue | No | ||
| contactId | No | ||
| anchorDate | Yes | ||
| customDays | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| maxOccurrences | No | ||
| templateDocumentId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses key behavioral traits: the template is snapshotted, not live-linked; no supply date is stored (each generated line gets its own period as Leistungsdatum); nothing is invoiced yet. It also notes autoIssue creates drafts unless on (P8), which is a significant behavior. This is excellent for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence but packs essential information. It is relatively concise for the complexity, though it could be better structured with bullet points for readability. The key details are front-loaded (purpose, then components, then behaviors).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (15 params, no output schema), the description does a good job covering core behaviors and key parameters. However, it misses some parameters (e.g., lines structure, idempotencyKey semantics) and doesn't mention return value or error scenarios. It's close to complete for a competent agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters. It does clarify 'interval months|quarterly|yearly|custom with customDays, anchorDate ISO YYYY-MM-DD', and mentions optional endDate, maxOccurrences, dueDays, autoIssue. However, it does not explain other parameters like lines, contactId, templateDocumentId beyond mentioning 'template (contact + positions, or templateDocumentId to snapshot an existing document)', which is partial. It adds value but not exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool creates a recurring invoice schedule ('Serienrechnung'), and specifies the key components (template, cadence, optional end conditions). It distinguishes from siblings like run_due_recurring by clarifying that nothing is invoiced yet, and the mention of snapshotting vs. live link adds specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly mentions run_due_recurring as the tool that materializes periods, which serves as a clear alternative. It also implies the appropriate context (creating a schedule) without explicitly listing exclusions, but it doesn't fully outline when not to use this tool (e.g., for one-off invoices).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_saved_viewA
Save a named filter, sort, column set and layout over an entity list. The view belongs to the calling session unless shared is true, which publishes it to everyone in the workspace and requires manage_saved_views.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| sort | No | ||
| layout | No | ||
| shared | No | ||
| columns | No | ||
| filters | No | ||
| isDefault | No | ||
| entityKind | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral disclosure. It conveys that the view is session-scoped unless shared, and that sharing publishes to the workspace and demands a specific permission. This is a meaningful side effect beyond a simple save, though it doesn't cover idempotency or conflict behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The first sentence immediately states the purpose and content, and the second adds the critical shared-scope behavior. Everything is front-loaded and necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (10 parameters, no annotations, no output schema), the description covers the main purpose and the shared edge case but omits details on idempotency, isDefault behavior, and return values. It is adequate for typical use but not fully comprehensive for all parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It maps the central parameters (name, filters, sort, columns, layout, shared) but leaves isDefault, idempotencyKey, entityKind, and workspaceId unexplained. For a 10-parameter tool this is partial but helpful semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool saves a saved view composed of a named filter, sort, column set, and layout over an entity list. It uses a specific verb ('Save') and resource, and the content details distinguish it from sibling operations like list_saved_views or update_saved_view.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides an implied usage context — you save a view when you want to capture a current list configuration. It includes a conditional requirement (shared requires manage_saved_views) but does not explicitly contrast with alternatives such as update_saved_view or list_saved_views, leaving selection logic partly implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_vendor_billA
Erfasse eine Kreditorenrechnung als Entwurf (a draft vendor bill: no posting, no ledger effect). amountMinor is integer Rappen and is the GROSS on the paper bill unless amountIsGross is false; taxCode is an input-side code (VST-M, VST-I, BEZUG, IMPORT) or absent for a vendor that charges no MWST. projectId tags the purchase to a B00 project for the B03 Projekterfolg (a reporting dimension like costCenterId: it prices nothing and appears on no journal leg). This is the agent path (P8): draft first, then post_vendor_bill as a separate, deliberate step.
| Name | Required | Description | Default |
|---|---|---|---|
| fxRate | No | ||
| dueDate | No | ||
| taxCode | No | ||
| billDate | Yes | ||
| currency | No | ||
| vendorId | Yes | ||
| projectId | No | ||
| receiptRef | No | ||
| supplyDate | No | ||
| amountMinor | Yes | ||
| workspaceId | Yes | ||
| costCenterId | No | ||
| amountIsGross | No | ||
| idempotencyKey | Yes | ||
| vendorReference | No | ||
| expenseAccountId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the critical behavioral trait: this creates a draft with no posting and no ledger effect. It also explains the semantics of projectId (a reporting dimension like costCenterId, prices nothing, appears on no journal leg) and amountMinor (integer Rappen, GROSS unless amountIsGross is false). This is substantial behavioral context beyond what the schema provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient, covering the draft state, amount semantics, tax code semantics, projectId semantics, and the workflow path in three sentences. It is front-loaded with the most important fact (draft, no posting). It loses one point because the German opening ('Erfasse eine Kreditorenrechnung als Entwurf') is immediately repeated in English, which is slightly redundant, and the sentence is long.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter tool with no output schema and no annotations, the description covers the highest-risk parameters (amountMinor, taxCode, projectId) and the draft-vs-post workflow. It does not explain the remaining 13 parameters (fxRate, dueDate, currency, receiptRef, supplyDate, costCenterId, idempotencyKey, vendorReference, expenseAccountId, workspaceId, vendorId, billDate, amountIsGross), but many are self-explanatory from their names. The critical workflow context (draft first, then post_vendor_bill) is present, which is the most important completeness factor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the most ambiguous parameters: amountMinor (integer Rappen, GROSS on paper bill unless amountIsGross is false), taxCode (input-side code with enumerated values VST-M, VST-I, BEZUG, IMPORT, or absent for no-MWST vendors), and projectId (reporting dimension, no pricing, no journal leg). This is exactly the semantic information an agent needs that the bare schema cannot convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Erfasse' = capture/create), a specific resource (Kreditorenrechnung = vendor bill), and the key state (as a draft/Entwurf with no posting, no ledger effect). It clearly distinguishes this from the sibling post_vendor_bill by explicitly naming it as a separate, deliberate step. The agent can tell exactly what this tool does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says this is the agent path (P8): draft first, then post_vendor_bill as a separate, deliberate step. It names the sibling tool post_vendor_bill as the alternative and gives the workflow context. It also explains when taxCode should be absent (vendor charges no MWST), which is a clear usage condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_workspaceC
Mint a new workspace (tenant) with its KMU chart of accounts.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| legalForm | No | ||
| baseCurrency | No | ||
| idempotencyKey | No | ||
| fiscalYearStart | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It reveals that a new tenant and its chart of accounts are created, but it does not mention required permissions, side effects, reversibility, provisioning scope, or what happens after creation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no filler or repetition. It is compact, though 'Mint' and 'KMU' are somewhat terse and jargon-like.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there is no output schema, no annotations, five undocumented parameters, and additionalProperties allowed, this description is not sufficient for an agent to invoke the tool reliably. It lacks parameter semantics, return/result expectations, and usage context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no parameter-level meaning. The agent receives no help understanding legalForm, baseCurrency, idempotencyKey, fiscalYearStart, or why name is required, despite the schema having no property descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Mint' a new workspace/tenant) and the resource being created, and adds meaningful scope by mentioning the KMU chart of accounts. It does not explicitly differentiate from closely related siblings like create_demo_workspace or bootstrap_workspace, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives such as create_demo_workspace, bootstrap_workspace, or onboard_client. There are no stated prerequisites, exclusions, or context cues beyond the general notion of creating a workspace.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
customer_balanceBRead-only
What one customer owes: their total open amount, the split across aging buckets in both the invoice currency and base currency, how many days their oldest overdue item has been outstanding, and any parked payment attached to them. onAccountMinor is reported positive for a Guthaben, so it comes out negative when what is parked is an unmatched refund: read the direction on the row rather than the sign of the total. An unknown customer reads as an empty balance rather than an error. Reads only.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| customerId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses two genuinely non-obvious behaviors: the onAccountMinor sign convention (positive for a Guthaben, negative for an unmatched refund, read the row direction not the total sign) and that an unknown customer returns an empty balance rather than an error. The final 'Reads only' is somewhat redundant with the annotation, but the added quirks provide real value beyond the structured metadata.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core return-value summary is front-loaded in the first sentence, with edge-case disclosures following in tight, purposeful sentences. The 'Guthaben' phrasing is slightly jargon-dense, but each sentence earns its place and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
In the absence of an output schema, the description does a strong job of telling the caller what it will get back and how to interpret the parked-payment row. However, with three undocumented parameters and no output schema, leaving the asOf semantics unexplained is a real gap for a tool whose output depends on an evaluation date; the description partially compensates for missing output richness but not for parameter coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the full burden for parameter documentation, and it mostly fails to deliver. customerId is only implicitly explained ('one customer implies the subject), but asOf is never mentioned at all — a surprising gap for a balance report where aging buckets are inherently date-sensitive — and workspaceId is not explained either. The description adds almost nothing about how the three parameters behave.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific subject ('What one customer owes') and enumerates the actual outputs: total open amount, aging buckets in two currencies, oldest-overdue days, and parked payments. This clearly identifies the resource and scope and distinguishes it implicitly from aggregate reports like aging_report. However, it never names a sibling tool or explicitly contrasts itself, so it stops short of the top score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use / when-not-to-use guidance and no alternative tools are named, despite very close siblings existing (aging_report, list_open_items, trial_balance). The intended use is inferable from 'what one customer owes,' but an agent choosing between a per-customer balance and an aging report gets no help from this description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dashboard_overviewARead-only
Die Übersicht (dashboard): one KPI tile wall for a date range, composed live from the read models the product already computes and never cached (pure read, F00 posts nothing and owns no tables). Core tiles: revenue | cash | ar_aging | ap_aging | utilisation | project_margin | mwst_due | stock_value, each an integer-Rappen (or basis-point) figure that equals its source verb's answer for the same filter: revenue from income_statement (the Nettoerlöse section, with a previous-window trendBp), cash from the A19 bank accounts' Kontoblatt closing balances (general_ledger), AR from aging_report, AP from list_vendor_bills, utilisation as the billable share of time_list minutes (B01 has no capacity model), margin from costing_pl_list, MWST due verbatim from vat_return for the settlement period containing the range end, stock value from stock_valuation_report under the books' method. Every tile embeds a drill descriptor (studioRoute, mcpTool, params) naming the EXISTING source tool behind the number: call that tool for the rows. Tiles the caller's A24 role may not see are omitted server-side into omitted[] (per-tile gates: read_books, read_sales, read_vat, read_master_data, time.read, costing.read); an unconfigured or unused module degrades its one tile to ok:false with the source's own code (needs_vat_config, needs_chart, needs_projects, needs_stock_items) while the rest still render; an empty workspace answers ok:true zero states, never an error. from > to answers invalid_range. savedViewId applies a G00 saved view (entityKind 'workspace', layout 'dashboard') to restrict and order tiles; an unresolvable view falls back to the full default with viewFallback:true and can never widen access.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | ||
| from | Yes | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes far beyond the readOnlyHint annotation: declares it 'never cached (pure read, F00 posts nothing and owns no tables),' details per-tile permission gating (omitted[] with per-tile gates such as read_books and time.read), and discloses degradation behavior (ok:false with codes like needs_vat_config), empty-workspace behavior, and from>to invalid_range semantics. Also discloses that savedViewId resolution 'can never widen access.' There is no contradiction with annotations — the description reinforces readOnlyHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core definition and scoping constraints, then enumerates tiles and behavioral edge cases roughly in order of importance. Every clause carries routing, gating, or error semantics; the tile-to-source enumeration is long but earns its place. The single dense paragraph is harder to scan at a glance than a structured list, which costs one point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and zero schema-level param descriptions, the description is remarkably complete: tile composition, units (integer-Rappen/basis points), source mapping, permission gating, degradation codes, empty state, invalid_range, and saved-view fallback are all covered in prose. Genuine gaps remain: the exact date string format for from/to and the top-level response envelope are not specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden, and it compensates well: from/to get date-range semantics plus the ordering constraint (from > to answers invalid_range), and savedViewId gets rich semantics (G00 saved view, entityKind 'workspace', layout 'dashboard', restriction/ordering, viewFallback:true on failure). workspaceId is only implicit (as the scope of A24 role gating and empty-workspace behavior), leaving a small gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens by identifying the resource precisely — 'one KPI tile wall for a date range' — and names its eight constituent tiles with their exact data sources (income_statement, aging_report, vat_return, stock_valuation_report, etc.). This distinguishes it from sibling source-report tools and from per-tile tools like dashboard_tile without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes drill-down behavior to the source tools: each tile 'embeds a drill descriptor (studioRoute, mcpTool, params) naming the EXISTING source tool behind the number: call that tool for the rows.' It also specifies when saved views apply and the fallback when a view cannot be resolved. It lacks an explicit when-not-to-use statement against the closest sibling (dashboard_tile), so exclusions are absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dashboard_tileARead-only
Eine Kachel der Übersicht in voller Tiefe: the same figure dashboard_overview shows for 'revenue' | 'cash' | 'ar_aging' | 'ap_aging' | 'utilisation' | 'project_margin' | 'mwst_due' | 'stock_value', plus its detail block (the aging buckets behind the Debitoren figure, the top and bottom projects behind the Marge, the MWST settlement period and filed flag, the per-bank-account closing balances) and the drill descriptor naming the source tool that owns the rows. The value is the source verb's own answer, recomputed on every call; F00 fabricates nothing. An unknown tile id answers unknown_tile naming the valid set; a tile the caller's A24 role may not see answers permission_denied outright (the same per-tile gate the overview applies as an omission); an unavailable module answers the tile with ok:false and the source's own structured code.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | ||
| from | Yes | ||
| tile | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description adds substantial behavioral context: it states the value is recomputed on every call, that F00 fabricates nothing (data integrity guarantee), and details three distinct error responses (unknown_tile, permission_denied, and ok:false with structured code for unavailable modules). This goes well beyond the annotation and fully discloses the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and somewhat rambling, with many parenthetical clauses and mixed language (German/English). While it is front-loaded with the core purpose, the additional detail on error cases and drill descriptors makes it longer than necessary. It is structured but could be more concise and better organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 parameters, no output schema, and no parameter descriptions, the description provides insufficient context. It explains the tile selection and error responses, but omits parameter semantics (especially from/to), the exact return structure, and how the drill descriptor is formatted. The description is not complete enough for an agent to correctly construct a valid request.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain the parameters. It only indirectly covers the 'tile' parameter (via 'unknown tile id' and the valid set) but does not explain 'workspaceId', 'from', or 'to'. There is no mention of date range semantics or workspace context. The description fails to compensate for the missing schema information for 3 of 4 parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: it returns a single dashboard tile in full depth, listing the exact tile types it supports (revenue, cash, ar_aging, etc.), the detail block, and the drill descriptor. It also distinguishes itself from dashboard_overview by explaining it shows the same figure but with additional depth, making its purpose specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description references dashboard_overview ('the same figure dashboard_overview shows') and clarifies that this is a drill-down for specific tiles. It implicitly communicates when to use this tool (when you need full tile detail) but does not explicitly state when not to use it or name alternatives beyond the implicit overview. It does describe error conditions, which helps the agent understand valid usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deals_createA
Lege einen Deal an: a sales opportunity on a C00 contact with an integer-Rappen value, born open in the first open stage of its pipeline (the first write seeds the default funnel Lead, Qualifiziert, Offerte, Gewonnen, Verloren when none exists). A non-base currency is converted ONCE at capture through the §H-FX resolver and frozen on the row (valueBaseMinor + fxRate, never client inputs); the creation lands on the contact timeline (OP5). A deal never posts: its value is an estimate, not a ledger row.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| stageId | No | ||
| currency | No | ||
| contactId | Yes | ||
| pipelineId | No | ||
| valueMinor | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No | ||
| expectedCloseOn | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does disclose important side effects: currency conversion is performed once via FX resolver and frozen on the row, the deal never posts to the ledger, and creation lands on the contact timeline. This is substantial transparency for a create operation. However, it does not mention whether the operation requires any permissions, what happens to an invalid currency, or what the response object is.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense block of German text with significant useful information. It front-loads the primary action and then packs in mechanics (pipeline seeding, FX conversion, no posting). While it is long, every clause adds behavioral context. It could benefit from sectioning or being slightly split into separate sentences for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create tool with 9 parameters and no output schema, this description covers the essential semantics: value type, currency handling, pipeline creation defaults, timeline side effect, and accounting implication. It uniquely clarifies that a deal never posts to the ledger and that FX is resolved at capture. The idempotencyKey and expectedCloseOn parameters are not mentioned, but the core behavior is thoroughly documented, which is unusual for a tool definition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does substantially. It explains valueMinor is an integer-Rappen value, and clarifies that currency conversion produces valueBaseMinor + fxRate, never client inputs. It does not explicitly document stageId, pipelineId, expectedCloseOn, or idempotencyKey, but it provides critical context on value and currency semantics that the bare schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Lege einen Deal an' (create a deal), specifies the resource (sales opportunity), and provides detailed scope: on a C00 contact, integer-Rappen value, born open in the first open stage, default funnel seeded when none exists. This distinguishes it from sibling tools like deals_update, deals_move, deals_mark, deals_to_quote.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use this tool: to create a sales opportunity on a contact, with explicit details about the lifecycle (born open, first open stage, default pipeline seeding). It does not explicitly name alternative tools for updating or moving deals, but given sibling context deals_update/deals_move exist, the description could have explicitly excluded them. The rules around currency conversion and no-posting behavior give clear context for when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deals_listARead-only
The pipeline board in one read (P5): every pipeline for the picker, the selected pipeline stages in sort order, its deals (open only by default; includeClosed or a status filter widens), each with weightedMinor = round-once(valueBaseMinor times probability over 100), and the weightedTotalMinor over open deals in the workspace base currency, so mixed-currency funnels sum in one currency (§H-FX). C03 consumes this read; C01 does not forecast. savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| contactId | No | ||
| pipelineId | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| includeClosed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide readOnlyHint=true; the description adds substantial behavioral detail: open deals are default, includeClosed or a status filter widens results, weightedMinor is computed as round-once(valueBaseMinor times probability / 100), and weightedTotalMinor aggregates open deals in the workspace base currency. It also explains savedViewId filter merging. These go well beyond the annotation and enrich the agent's understanding of side effects and computations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but purposeful: it front-loads the core purpose ('pipeline board in one read'), then layers filters, calculations, currency logic, and saved view behavior. The use of internal cross-references (P5, C03, C01, G00, §H-FX) adds precision for those who know the spec but could be noise; nonetheless, every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with 6 parameters and no output schema, the description covers the essential output facets (pipelines, stages, deals, computed values, totals) and filter semantics. It does not mention pagination, sorting, or error conditions, but for the stated purpose it is sufficiently complete. The internal references may reduce universal clarity but do not leave critical gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description carries the full burden. It adds meaning for includeClosed and status ('widens'), savedViewId ('applies a saved view... merged underneath'), and implies pipelineId via 'selected pipeline stages' and workspaceId via 'workspace base currency'. However, contactId is not mentioned, and some parameters like workspaceId/pipelineId are only indirectly covered. Still, it compensates well for most parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'The pipeline board in one read' and enumerates exactly what is returned: all pipelines, selected pipeline stages in sort order, deals with defaults and widened filters, computed weightedMinor, and weightedTotalMinor. This clearly distinguishes it from sibling deal tools like deals_create, deals_update, and deals_move, and from forecasting tools like forecast_weighted_pipeline.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions 'C03 consumes this read; C01 does not forecast,' which hints at internal consumers but does not explicitly tell an AI agent when to choose this tool over alternatives such as forecast_weighted_pipeline or a simpler deals list. There is no direct 'use this when...' or 'instead of...' guidance, so the usage context is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deals_log_activityA
Erfasse eine Aktivität auf einem Deal (note, call, email, meeting or task, the single OP5 kind enum): appends to the contact timeline through contacts_log_activity with the dealId stamped, newest-first on read. An optional reminderAt additionally mints one linked E03 task (entityKind deal) through tasks_create, so its assignee polls tasks_reminders_due like every other follow-up; a past reminder refuses with E03s own reminder_in_past. One idempotency key covers note + task together.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | ||
| kind | Yes | ||
| dealId | Yes | ||
| occurredAt | No | ||
| reminderAt | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the side effect of appending to the contact timeline, the optional creation of a linked task via tasks_create, the 'newest-first on read' ordering, the rejection of past reminders, and the scope of the idempotency key covering both note and task. It does not cover all potential side effects (e.g., deal modification) but offers substantial behavioral detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph with multiple clauses and internal references (OP5, E03), making it harder to parse quickly. It includes valuable information but is not well-structured or front-loaded. It is concise in length but not optimally organized for agent consumption.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (7 parameters, optional task creation, idempotency, error handling) and no output schema or annotations, the description covers the main behavioral aspects but omits details like return value, potential errors beyond the past-reminder case, and how the deal timeline itself is affected. It is adequate for basic invocation but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains the 'kind' enum values, the behavior of reminderAt (including the error case), and how idempotencyKey applies to both note and task. It does not explicitly cover workspaceId, dealId, body, or occurredAt, though their meanings are reasonably inferable from context. Partial compensation for missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool records an activity on a deal, listing the allowed kinds (note, call, email, meeting, task) and explaining it appends to the contact timeline via contacts_log_activity with the dealId stamped. It references the underlying mechanism and distinguishes itself from plain contact logging by adding deal context, though it does not explicitly contrast with sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context by explaining the delegation to contacts_log_activity and tasks_createhare and mentioning that reminders work like other follow-ups via tasks_reminders_due. However, it does not explicitly state when to prefer this over alternatives or provide exclusions, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deals_markA
Schliesse einen Deal ab oder öffne ihn wieder: the ONE door to status. won moves the deal into the pipeline outcome stage at probability 100; lost requires a lostReason (lost_reason_required otherwise) and lands at 0; open reopens a misclicked deal into the first open stage. A no-change re-mark is a state assertion and emits nothing; a genuine close emits deal.won or deal.lost. Every terminal move lands on the contact timeline (OP5).
| Name | Required | Description | Default |
|---|---|---|---|
| dealId | Yes | ||
| status | Yes | ||
| lostReason | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations to carry the burden, the description discloses important behavior: lost without lostReason errors as lost_reason_required, no-change re-marks emit nothing, genuine closes emit deal.won/deal.lost, and terminal moves land on the contact timeline. This is rich, non-obvious behavior beyond the bare schema, adding real value an agent would otherwise miss.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, information-dense, and front-loaded with the core purpose. Every sentence adds a distinct behavioral fact, from status transitions to event emission to timeline side effects, with no filler or repetition of schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description covers the essential context: status semantics, error condition, event emission, reopening behavior, and timeline side effects. An agent has enough to invoke the tool correctly and predict its effects for all three status values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description compensates strongly for status and lostReason: it defines the allowed statuses, the lostReason conditionality, and the consequence of omitting it. It leaves workspaceId, dealId, and idempotencyKey unexplained, but these are self-evident and the core parameter semantics that matter are clarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear action: close a deal or reopen it, and then enumerates the exact statuses it handles: won moves to outcome stage, lost closes at 0, open reopens to the first open stage. It identifies the resource (deal) and the specific verb 'mark status', and positions itself as the single tool for deal status changes, distinguishing it from stage-movement siblings like deals_move.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete guidance for each status value: won closes, lost requires lostReason, open reopens a misclicked deal, and a no-change re-mark is a state assertion. It does not explicitly name sibling alternatives or state when not to use this tool, so it stops short of full when-to-use versus alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deals_moveA
Verschiebe einen Deal in eine andere Phase of its own pipeline: updates the stage, re-defaults the probability to the target stage unless a hand pinned it, and logs the move on the contact timeline (OP5). A stage outside the deal pipeline refuses with stage_not_in_pipeline; a stage flagged won/lost refuses with terminal_stage_use_mark, because deals_mark is the ONE door to a terminal status. Emits deal.stage_changed.
| Name | Required | Description | Default |
|---|---|---|---|
| dealId | Yes | ||
| stageId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses side effects (updates stage, resets probability unless pinned, logs on timeline, emits deal.stage_changed) and error conditions with specific error codes. This gives the agent a complete picture of the tool's behavior and consequences.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence but each clause adds value: action, probability behavior, timeline logging, error conditions, and event emission. It is front-loaded with the primary action and not overly verbose. However, the run-on structure with semicolons could be clearer, so not a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers core behavior, error cases, and the emitted event. It does not explain the idempotencyKey parameter or the return value, but with no output schema, these are secondary. The inclusion of specific error codes and the deals_mark alternative makes it fairly complete for an agent to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter documentation. However, it does not explain any of the four parameters (workspaceId, dealId, stageId, idempotencyKey). It only implies stageId via 'target stage' and dealId via 'Deal', but provides no details on format, meaning, or constraints. This is a critical gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the operation: moving a deal to another phase in its own pipeline. It specifies the verb (Verschiebe), the resource (Deal), and the scope (own pipeline). It also differentiates from the sibling deals_mark by explicitly stating that terminal stages require deals_mark, making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when not to use this tool: for stages outside the pipeline (refuses with stage_not_in_pipeline) and for terminal stages (refuses with terminal_stage_use_mark). It further directs the agent to the correct alternative (deals_mark) for terminal status changes, providing clear routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deals_to_quoteB
Wandle einen Deal in eine Offerte um: delegation only (US-C01.5). Invokes the registered quote-creation verb (create_document, type quote, until C02 rebinds the seam) through the shared dispatch with the deal contact and value as the single seed line, stores the returned quote id on the deal, and logs the conversion (OP5). Idempotent on the deal: one carrying a quoteId answers it and spawns nothing. The quote verb own capability applies to the caller; C01 never bypasses it.
| Name | Required | Description | Default |
|---|---|---|---|
| dealId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it discloses substantial behavior: it is delegation-only, invokes create_document via shared dispatch, stores the returned quote id on the deal, and logs the conversion (OP5). The idempotency guarantee and the capability note ('The quote verb own capability applies to the caller; C01 never bypasses it') clarify authorization and repeat-call safety. The opaque internal codenames slightly erode transparency, but the key behaviors are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and the purpose statement is front-loaded, but it is one dense paragraph packed with opaque internal references (US-C01.5, C02, OP5) that an agent cannot resolve. Every sentence does add behavioral information, yet the jargon reduces the value density in an otherwise short definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param tool with no annotations, no output schema, and 0% schema coverage, the description explains the core behavior and side effects well but leaves gaps: the return value is never described (it only mentions storing the quote id on the deal), and the parameters are not documented. The unexplained reference codes further reduce completeness for an agent that must call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explicitly define dealId, workspaceId, or idempotencyKey. It only indirectly implies semantics — 'the deal contact and value as the single seed line' hints at what deal data is used, and the idempotency clause loosely maps to the idempotencyKey parameter. Given the 0% coverage, the description fails to fully compensate for the missing parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The German opening 'Wandle einen Deal in eine Offerte um' states the verb and resource clearly: convert a deal into a quote. It further specifies the mechanism — delegation to the registered create_document verb with type quote — which distinguishes it from direct document-creation siblings, though it never names an alternative explicitly. The internal codes (US-C01.5, C02, OP5) add noise but do not obscure the core purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies its usage context (converting a deal into a quote) and adds an important idempotency caveat: a deal already carrying a quoteId will answer and spawn nothing. However, it never explicitly states when to prefer this tool over siblings such as quotes_create, quotes_convert, or convert_document, nor does it name any exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deals_updateA
Bearbeite einen offenen Deal from a patch: retitle, re-point the contact, reschedule the expected close, or pin a manual probability (which stage moves then stop re-defaulting). A patch naming valueMinor or currency re-runs the §H-FX capture and re-freezes valueBaseMinor + fxRate; any other patch leaves the frozen trio byte-identical. A closed deal refuses with deal_closed: reopen it through deals_mark first.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | Yes | ||
| dealId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so thoroughly. It discloses that patches containing valueMinor or currency trigger a re-run of the H-FX capture and re-freeze valueBaseMinor + fxRate, while other patches leave the 'frozen trio' byte-identical. It also reveals the stage-move side effect of pinning probability and the error behavior for closed deals (deal_closed).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise—four sentences with no padding—but the mixing of German and English ('Bearbeite ... from a patch') and dense jargon (H-FX capture, frozen trio, byte-identical) slightly reduces readability. It is still efficient and front-loaded with purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the key behavioral requirements for a mutation tool with no annotations or output schema. However, it does not clarify whether patches are atomic, the exact meaning of 'stage moves then stop re-defaulting' is vague, and the idempotencyKey parameter is not explained. These are gaps, though the essential semantics are well covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must add meaning to the parameters. It explains the effect of each patch field: retitle, re-point contact, reschedule expected close, pin manual probability (which moves stage), and that valueMinor/currency re-run FX capture. This goes far beyond the bare schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the operation ('Bearbeite einen offenen Deal' = edit an open deal) and enumerates the specific fields that can be patched (title, contact, expected close, probability, valueMinor, currency). The 'open deal' condition and 'closed deal refuses' clearly distinguish it from deal lifecycle tools like deals_create or deals_mark.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit when-not guidance: a closed deal refuses and must first be reopened via deals_mark. The 'open deal' precondition is stated upfront, and the description names the exact alternative (deals_mark) for the refusal case, leaving no ambiguity about when to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
define_fieldA
Define a typed custom field on any registered entity, or patch an existing one by its key. The key may not shadow a column the entity already has, and once any value is stored the type is frozen (changing it would silently corrupt what is already there). A field defined by the agent actor lands as a draft and is invisible until confirm_field releases it. CONSEQUENCE: Adds a workspace-visible custom field to every record of the chosen kind.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| sort | No | ||
| type | Yes | ||
| options | No | ||
| required | No | ||
| labelI18n | Yes | ||
| entityKind | Yes | ||
| workspaceId | Yes | ||
| defaultValue | No | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses key constraints: the key may not shadow a column, the type is frozen once values are stored, and agent-defined fields land as invisible drafts until confirmed. It also states the consequence of adding a field to every record. This is rich behavioral context, though it does not mention idempotency or error conditions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph with each sentence adding value: purpose, constraints, draft workflow, and consequence. It is somewhat dense but not wasteful; the CONSEQUENCE label clearly highlights the side effect.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 10 parameters (5 required) with nested objects and no output schema, the description is incomplete. It does not explain how to structure labelI18n or options, and does not describe the response. The draft and freezing behavior is useful but insufficient for an agent to call this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explain any parameters beyond implying 'key' as an identifier and 'type' as the field type. It does not clarify labelI18n, options, required, defaultValue, sort, entityKind, workspaceId, or idempotencyKey. An agent would struggle to construct correct arguments without additional information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool defines a typed custom field on any registered entity or patches an existing one by key. It includes a CONSEQUENCE note describing the workspace-visible effect, which distinguishes it from operations like archiving or setting values.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the draft workflow and mentions confirm_field for releasing fields, but it does not explicitly name alternative tools (e.g., list_field_defs to check existing fields, set_field_value to set values) or state when not to use this tool. The context is helpful but lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
define_roleA
Create a custom role, or reshape one of the three editable built-ins, as a named subset of the capability registry. A name outside the registry is refused (unknown_capability): a role can only ever recombine capabilities that already gate a real verb. CONSEQUENCE: Redefines which capabilities a role holds, for every member who has that role.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| roleId | No | ||
| workspaceId | Yes | ||
| capabilities | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses that the tool mutates role definitions and explicitly warns: 'CONSEQUENCE: Redefines which capabilities a role holds, for every member who has that role.' It also reveals a validation rule (refuses names outside the registry). This goes beyond a simple 'create' verb and provides meaningful side-effect awareness, though it does not mention any authorization requirements or reversibility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact—three sentences total—and front-loads the core purpose, then adds the validation constraint and the consequence. Each sentence earns its place, with no filler. The structure is logical and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with side effects and no output schema, the description is incomplete. It covers the high-level behavior and a key consequence, but omits essential details about how to specify the target role (name vs roleId), the format of capabilities, the role of workspaceId, and the idempotencyKey's purpose. An agent would need to infer too much about the parameters to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, so the description must explain parameters. It explains 'name' and 'capabilities' (as a named subset of the registry), but does not clarify 'workspaceId', 'roleId', or 'idempotencyKey'. The roleId is only indirectly implied via 'reshape one of the three editable built-ins', and workspaceId and idempotencyKey are left entirely unexplained. This leaves the agent guessing about required context and idempotency semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: create a custom role or reshape a built-in role, defined as a subset of the capability registry. It distinguishes itself from role-related siblings like list_roles, set_role, and archive_role by focusing on defining/redefining capabilities, not listing or assigning roles. The mention of 'custom role' and 'three editable built-ins' adds specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool (when you need to create or modify role definitions) but does not explicitly state when not to use it or name alternatives. For example, it doesn't say 'use list_roles to view roles' or 'use set_role to assign roles to members'. The context is clear, but exclusions and alternative routing are absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_accountA
Hard-delete an account that never carried a posting.
| Name | Required | Description | Default |
|---|---|---|---|
| accountId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. 'Hard-delete' clearly indicates permanent, destructive action, and the posting condition suggests a safety guard. However, it does not disclose what happens if the account has postings (e.g., error), whether the deletion is reversible, or permission requirements. It provides some behavioral context but leaves operational specifics unspecified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that front-loads the critical action and constraint. Every word contributes meaning; there is no filler or redundancy. It is an ideal example of concise, effective tool documentation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive operation with no annotations and no output schema, the description provides the core behavior and a key precondition, but omits important contextual details such as error handling if the condition is violated, irreversibility beyond the term 'hard-delete,' and any guidance on the idempotencyKey parameter. It covers the essentials but is not fully complete for an agent to invoke it without ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it does not explain any of the three parameters (workspaceId, accountId, idempotencyKey). It implicitly refers to the account being deleted via the word 'account,' but does not clarify which parameter identifies it or how idempotencyKey is used. This leaves parameter meaning almost entirely to the schema, which itself offers no descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the exact action ('Hard-delete') and the specific resource ('account'), and adds a clear distinguishing condition ('that never carried a posting'). This separates it from archive_account, unarchive_account, and other account-related operations, making the tool's purpose immediately identifiable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The condition 'that never carried a posting' explicitly communicates when it is appropriate to use this tool, implying that accounts with postings should not be hard-deleted(likely archived instead). While it doesn't name an alternative tool, it gives a clear safety criterion for usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_backupB
Delete a local backup/export artifact and its history row. Storage housekeeping only: a backup file carries no legal retention lock of its own.
| Name | Required | Description | Default |
|---|---|---|---|
| backupId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses that it deletes both the artifact and its history row, and clarifies that there is no legal retention lock. However, it does not mention side effects like irreversibility, required permissions, or the behavior of the idempotencyKey. With no annotations provided, the description carries the full burden and falls short on these aspects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise—two sentences—and front-loaded with the primary action. The second sentence adds valuable context without fluff. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has three required parameters, no output schema, and no parameter descriptions. The description does not explain return values, the role of idempotencyKey, or any prerequisites or consequences. For a deletion operation, this is notably incomplete; an agent would need more detail to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, and the description provides no meaning for the three parameters (workspaceId, backupId, idempotencyKey). It does not compensate for the lack of schema descriptions, leaving the agent to infer their purposes from names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (delete) and the specific resource (local backup/export artifact and its history row), which distinguishes it from sibling tools like create_backup, restore_backup, or verify_backup. It also adds context that this is storage housekeeping, further clarifying its scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'storage housekeeping only' implies this is for cleanup, but it does not explicitly state when to use it versus alternatives, nor does it mention exclusions (e.g., when not to use it, or that restore_backup should be used for recovery). The guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_cost_centerA
Hard-delete a cost centre that no posted line references.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| costCenterId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. 'Hard-delete' indicates permanent, irreversible removal, which is a strong signal. However, it does not disclose failure behavior (e.g., what happens if references exist) or any other side effects. The permanence is implied but not elaborated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no wasted words. It is front-loaded with the action and constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no annotations and no output schema, this is insufficient. It does not differentiate from archive_cost_center (soft delete), does not explain failure modes, and does not mention whether deletion is reversible or requires specific permissions. The agent needs more context to make an informed choice.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no information about the parameters (workspaceId, costCenterId). The names are self-explanatory, but the description adds zero value beyond the schema, leaving the agent with no extra guidance on formats, types, or relationships.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Hard-delete') and resource ('cost centre') with a clear constraint ('no posted line references'). It is immediately distinguishable from create_cost_center, archive_cost_center, and unarchive_cost_center.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the condition for safe use ('no posted line references'), implying this is only appropriate when no references exist. It does not explicitly mention alternatives like archive_cost_center, but the condition provides clear context for when deletion is allowed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_draftB
Delete a draft journal entry (never a posted one).
| Name | Required | Description | Default |
|---|---|---|---|
| entryId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the destructive nature ('Delete') and the restriction to drafts, which is useful. But it does not mention irreversibility, idempotency implications, or any side effects, so it is only partially transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with zero fluff; the core action is front-loaded and the critical constraint is appended. This is an appropriately sized description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive operation with three required parameters and no output schema, the description omits critical information: what happens upon deletion (irreversibility), behavior on non-existent drafts, idempotency guarantees, and required permissions. It is far from complete for an agent to safely invoke.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description provides no explanation for workspaceId, entryId, or idempotencyKey. The agent gets no guidance on parameter meaning, format, or relationships, making this a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Delete') and resource ('draft journal entry') with an explicit qualifier ('never a posted one'), which clearly distinguishes it from post_entry, reverse_entry, and other journal operations. The safety constraint is meaningful and disambiguates from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is only for drafts through the qualifier 'never a posted one', which acts as a when-not condition. However, it does not explicitly name alternative tools (e.g., reverse_entry, post_entry) or explain when to prefer them, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_itemA
Hard-delete an item that is referenced nowhere (no document line, price-list row, variant, or stock movement). A referenced item is refused with item_referenced and its reference kinds; archive it instead so posted history stays resolvable.
| Name | Required | Description | Default |
|---|---|---|---|
| itemId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of behavioral disclosure. It clearly states that the deletion is irreversible ('hard-delete'), that referenced items are refused with item_referenced, and why archiving is preferable. It does not mention consequences like audit trail or reversibility beyond 'hard-delete', but the core destructive nature and refusal behavior are transparent enough.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no redundancy, and the main action is front-loaded. The first sentence states precisely what the tool does, and the second gives the refusal behavior and the alternative. Everything earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a delete tool with no output schema and no annotations, the description covers the essential context: when to use, what refusal looks like, and the recommended alternative. It omits parameter details (already noted) but otherwise gives an agent enough to decide and invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the parameters. However, it provides no information about itemId, workspaceId, or idempotencyKey. While the names are suggestive, the description adds no meaning beyond the schema, and for a delete operation, the idempotencyKey could benefit from clarification. This is a clear gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the action ('Hard-delete an item') and the resource, and differentiates it from sibling archive_item by specifying that it only applies to items referenced nowhere. It names the refusal condition and gives a reason to archive instead, making the tool's purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage guidance: use this tool only for unreferenced items, and if the item is referenced, it will be refused—use archive_item instead. This clearly tells an agent when to call this tool and when to choose the alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_saved_viewA
Delete a saved view outright. A true delete rather than an archive, because a view mints no state and holds no history; deleting a shared one requires manage_saved_views.
| Name | Required | Description | Default |
|---|---|---|---|
| viewId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure. It clearly states this is a 'true delete' with no retained state or history, and flags the permission needed for shared views. It does not explicitly mention irreversibility, but 'true delete' strongly implies that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the primary action, and every clause adds valuable context – the distinction from archive, the reason (no state/history), and the permission caveat. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description gives the essential behavioral model and one permission constraint, but it omits parameter semantics, idempotency behavior, and return/error information. For a simple delete operation this is on the edge of adequacy but still leaves gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for explaining parameters. It does not mention workspaceId, viewId, or idempotencyKey at all. The parameter names are somewhat self-explanatory, but the purpose and required usage of idempotencyKey, and how workspaceId scopes the deletion, remain undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Delete a saved view outright.' It also distinguishes itself from an archive operation, making its purpose unmistakable even among many sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the key usage context – it is a permanent delete, not an archive – and specifies a permission requirement for shared views ('deleting a shared one requires manage_saved_views'). It does not enumerate alternatives, but there is no archive_saved_view sibling, so this is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delivery_note_createA
Erstelle einen Lieferschein: drafts a delivery note for a confirmed orders stock-tracked lines from one D01 location. With no explicit lines it defaults to each lines undelivered qty; an explicit line may ship less (a partial). Issues nothing and moves no stock. Over-delivery (qty over the outstanding remainder) is refused with over_delivery naming the outstanding qty; a zero/fractional qty is refused (invalid_qty).
| Name | Required | Description | Default |
|---|---|---|---|
| lines | No | ||
| locationId | Yes | ||
| workspaceId | Yes | ||
| salesOrderId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses no side effects, default quantity behavior, partial shipment support, and specific error conditions (over_delivery, invalid_qty). It does not cover idempotency behavior, but the core side effects and rejection rules are clearly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with no filler: the core action is front-loaded, followed by default behavior, then error constraints. Every sentence adds necessary decision-making information for the agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is notably complete: it covers prerequisites, defaulting, partials, error conditions, and absence of side effects. It does not describe the return value or explicitly point to delivery_note_issue as the next step, but the tool is still well specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description compensates strongly for the complex 'lines' parameter by explaining defaults, partials, and validation failures. It also clarifies the location restriction to a single D01 location. Other parameters like workspaceId and salesOrderId are self-explanatory, though idempotencyKey semantics are not described.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'drafts a delivery note' for a confirmed order's stock-tracked lines from one D01 location. It explicitly distinguishes itself from related actions by stating 'Issues nothing and moves no stock', which differentiates it from delivery_note_issue and stock_move.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear context for when to use the tool: on confirmed orders, stock-tracked lines, and a single D01 location. The statement 'Issues nothing and moves no stock' implies when not to use it, but it does not explicitly name alternative tools or provide an explicit when-not-to-use directive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delivery_note_issueA
Liefere aus (P8, stock leaves the shelf): issues a drafted delivery note. Mints ONE D01 issue movement per line through stock.move (OP2, D03 never writes stock_movement itself), stores each stock_movement_id, raises delivered_qty and advances the order (partially_delivered / delivered). Atomic and idempotent: over-delivery and insufficient stock are pre-checked (a refusal writes zero rows), a D01 refusal rolls the whole note back, and a retry re-issues nothing. carries actor for the audit trail.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| workspaceId | Yes | ||
| deliveryNoteId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it is exceptionally transparent. It discloses side effects (stock leaves shelf, delivered_qty raised, order advances), internal mechanics (one D01 movement per line via stock.move), and failure semantics (pre-checks, zero-row refusal, full rollback on D01 refusal, retry re-issues nothing).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then gives dense, relevant operational details. It is longer than average, but the atomicity, idempotency, and rollback behavior are material and each clause adds value; the heavy use of jargon is the main cost.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating stock operation with no annotations and no output schema, the description is nearly complete: it covers inputs, side effects, state transitions, failure behavior, and audit context. It lacks only an explicit statement of what the tool returns and how idempotencyKey should be supplied by the caller.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning for 'actor' by saying it is carried for the audit trail and hints at idempotency via 'retry re-issues nothing,' but it does not explicitly map idempotencyKey usage or explain the workspace/delivery note parameters beyond their obvious names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Liefere aus' and 'issues a drafted delivery note,' making the action unambiguous. It also distinguishes itself from siblings like delivery_note_create and delivery_note_render by specifying that this is the issuing step where stock leaves the shelf, so an agent can tell it apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: on a drafted delivery note, at stage P8, when stock leaves the shelf. It does not explicitly name alternatives or state when not to use it, but the 'drafted' qualifier and side-effect framing provide adequate usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delivery_note_renderA
Drucke / lade den Lieferschein (local artifact, OP4): renders the Lieferschein PDF for an ISSUED note, files it in E00 (entity_kind delivery_note) and locks OR 958f retention off the shipment date. The PDF carries NO VAT statement (a delivery note is a Beleg, not a taxable document; MWST arises on the A11 invoice). Idempotent: a re-render returns the existing E00 document, one Beleg per note. A draft note is refused (not_issued). The OSS core renders and files, then STOPS; emailing it is cloud-tier.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| deliveryNoteId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries the behavioral burden and does so impressively. It discloses idempotency, permanent filing in E00, a retention lock based on shipment date, the absence of a VAT statement, and the one-Beleg-per-note invariant, plus the refusal behavior for drafts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but not wasteful; the core purpose is front-loaded and the additional clauses each add meaningful context. Some jargon (E00, OR 958f, OSS core) is compact but domain-specific, and the structure could be slightly clearer for an agent unfamiliar with these internal codes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with no annotations and no output schema, the description covers an unusually complete amount: prerequisites, idempotency, filing behavior, retention implications, VAT semantics, and tier boundaries. It does not describe return values or error response shapes, which would be needed for fully complete operational guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it never explains workspaceId, deliveryNoteId, or idempotencyKey as parameters. It implies idempotency behavior, which hints at idempotencyKey, but does not explicitly map parameters to their purpose or required format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the exact action with a specific verb and resource: renders the Lieferschein PDF for an ISSUED note, then files it in E00 as a delivery_note entity. It also clearly differentiates from related operations by stating that draft notes are refused and that emailing is out of scope for the OSS core.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context: only issued notes are valid, drafts are refused with not_issued, and the tool stops after rendering/filing, with emailing happening in the cloud tier. It does not explicitly name sibling alternatives like delivery_note_issue or delivery_note_create, so the exclusion is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delivery_statusARead-only
Inspect the running TILL delivery process: mode (up | mcp | serve | agent_session), version, the schema generation of the open database, the path of the open database file (dbPath, ":memory:" on an ephemeral run), the bound loopback host and port, whether the built Studio is served, and the local scheduler (enabled, last tick, next tick). Pre-workspace: it describes the process, not any tenant, so it takes no workspaceId. In agent_session mode your books run inside the agent runtime and residency is the runtime vendor, not your machine (M00 US-M00.6).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds valuable behavioral context beyond that: it details the fields reported, explicitly notes that it takes no workspaceId, and explains the agent_session residency nuance. It does not mention potential absence of a database or failure modes, but the read-only nature and field list provide strong transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a dense but well-organized single paragraph of three sentences. It front-loads the primary purpose and enum values, then lists the inspected fields, and finally adds usage caveats. It is appropriately sized for a diagnostic tool with many reported fields, though slightly verbose with the M00 reference.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, the description fully enumerates all expected return fields (mode, version, schema generation, dbPath, host/port, Studio served, scheduler details). It also covers the pre-workspace context and agent_session residency, making the tool's behavior clear enough for an agent to invoke it correctly without guessing the result shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the baseline for zero parameters is 4. The description adds clarifications about what the tool does not take (workspaceId) and pre-workspace usage, which is useful guidance. With 100% schema coverage (empty schema), there is nothing else to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Inspect') and resource ('running TILL delivery process'), and enumerates the exact data points returned (mode, version, schema generation, dbPath, host/port, Studio served, scheduler). It clearly distinguishes this from tenant-level status tools by emphasizing it describes the process, not a tenant, and takes no workspaceId.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use it: it works pre-workspace and describes the process, so no workspaceId is needed. It also clarifies behavior in agent_session mode. However, it does not explicitly name alternative sibling tools (e.g., runtime_status, env_status) or state when not to use it in favor of those, so it falls short of the highest bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
describe_rate_feedARead-only
Report which published rate series MWSTV Art. 45 admits, the endpoint each one lives at, how its validity dates and per-unit quotations work, and the pages that establish it. Call this first: the engine performs NO network I/O, so a caller fetches the endpoint itself and hands the response body to import_exchange_rates.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Even though the `readOnlyHint` annotation already signals safety, the description adds crucial behavioral detail: 'the engine performs NO network I/O' and that the caller must fetch the endpoint independently. This goes beyond the annotation and helps the agent correctly orchestrate the call, avoiding false assumptions about network side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero fluff. The first sentence states the core purpose and the second gives the critical usage instruction. Information is front-loaded and every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple schema (one param), minimal annotations, and no output schema, the description fully covers what the tool returns (series, endpoints, validity, quotations, pages) and how to use it (call first, pass response to `import_exchange_rates`). Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has a single required `workspaceId` with no description, and schema description coverage is 0%. The description does not mention `workspaceId` at all, leaving its purpose only implicitly tied to the workspace context of rate feeds. With such low coverage, the description should have compensated by clarifying the parameter, but it doesn't.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a specific verb and resource: 'Report which published rate series MWSTV Art. 45 admits'. It precisely enumerates what the tool covers (endpoint, validity dates, per-unit quotations, establishing pages), making its function unmistakable. It also distinguishes itself from the related `import_exchange_rates` by explicitly marking itself as the preceding step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Call this first' and explains the exact workflow: 'a caller fetches the endpoint itself and hands the response body to `import_exchange_rates`.' This gives unambiguous sequencing and rationale (the engine performs no network I/O), leaving no doubt about when and how to use the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
detect_anomaliesARead-only
Surface likely mistakes as flags, never auto-corrections: probable duplicate entries, stale drafts, missing tax codes and round-number outliers over the posted ledger (optional period YYYY or YYYY-MM). Writes nothing; the user confirms before any fix is drafted.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description adds 'Writes nothing; the user confirms before any fix is drafted' and 'never auto-corrections.' This gives the agent important behavioral details about side effects and the interaction flow, adding value over the annotation's simple read-only flag.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The core action ('Surface likely mistakes as flags') is front-loaded, followed by the specific anomaly types and the safety constraint ('Writes nothing'). Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 parameters, no output schema) and the readOnlyHint annotation, the description fully covers what an agent needs: the action, the scope, the optional period format, and the confirmation behavior. Nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates partially. It explains the period parameter's format ('optional period YYYY or YYYY-MM') but does not explicitly describe workspaceId. However, workspaceId is self-explanatory as a workspace identifier and the description implies its role via 'over the posted ledger.' This adds meaning beyond the raw schema, though it could be more explicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Surface likely mistakes as flags, never auto-corrections' and lists specific anomaly types (duplicate entries, stale drafts, missing tax codes, round-number outliers) over the posted ledger. It distinguishes itself from sibling tools by emphasizing it never auto-corrects and requires user confirmation, making its scope unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives context on when to use it ('over the posted ledger') and mentions an optional period (YYYY or YYYY-MM). However, it does not explicitly list alternatives or conditions for when not to use it, which would earn a 5. It provides clear context but no exclusions or named sibling comparisons.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
disable_automation_ruleA
Stop a rule firing, immediately. Deliberately requires NO capability at all: a stop button that needs a permission is not a stop button, and disabling only ever prevents a write, never causes one.
| Name | Required | Description | Default |
|---|---|---|---|
| ruleId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It clearly discloses the important behavioral trait that the tool deliberately requires no capability and 'only ever prevents a write, never causes one.' It omits idempotency and what happens if the rule is already disabled, but the core behavioral safety profile is well explained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short, front-loaded, and every sentence contributes either the core action or the deliberate capability/side-effect profile. The rationale sentence is slightly rhetorical but still earns its place by emphasizing the no-permission design decision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter tool with no output schema or annotations, the description covers the main behavioral context well. Missing context includes reversibility, idempotency, and whether other tools should be used for permanent removal. The description is adequate but has clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does not explain what workspaceId and ruleId mean or how they relate to the stopped rule. The parameter names are self-descriptive, but the description adds no parameter-level semantics beyond the bare schema properties.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the exact action with 'Stop a rule firing, immediately,' which clearly distinguishes it from sibling tools like enable_automation_rule, archive_automation_rule, or update_automation_rule. It also underscores the disabling nature by noting it 'prevents a write' rather than causing one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context: use this as an immediate stop button that requires no capability. However, it never explicitly states when to choose this over alternatives like archive_automation_rule or enable_automation_rule, and it does not provide when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
disable_pluginA
Deaktiviere eine Erweiterung ohne sie zu deinstallieren (US-G02.3): sweep every capability registration the plugin holds (its MCP tools vanish, its Studio screen unmounts, its report source and automation actions stop resolving), stop its sandbox, and flip status to disabled, keeping the manifest and granted permissions so a later enable_plugin needs no re-install. History (audit rows, past automation runs) is untouched.
| Name | Required | Description | Default |
|---|---|---|---|
| pluginId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it lists all side effects (capability registrations removed, sandbox stopped, status flipped, history preserved) and explains what is kept. This is thorough and transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence that front-loads the main action and then details effects. It's not overly verbose but could be more structured; every clause adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains the tool's effects thoroughly but omits any mention of return value or error conditions, and doesn't clarify the idempotencyKey parameter. Given the tool's complexity and lack of output schema, this is a notable gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description provides no explanation of pluginId, workspaceId, or idempotencyKey. While the first two are self-explanatory, idempotencyKey's purpose is not addressed, so the description fails to compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool deactivates a plugin without uninstalling it, explicitly distinguishing it from uninstall_plugin and noting that later enable_plugin needs no re-install. It enumerates specific effects (MCP tools vanish, screen unmounts, etc.), making the purpose unmistakable and differentiating from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when a temporary deactivation is needed, keeping manifest and permissions for later enable. It contrasts with uninstallation but doesn't explicitly name alternatives or when-not-to-use, though the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
discard_demo_workspaceA
Hard-delete a demo workspace and every row under it, because nothing in a demo is a real fiscal record. Refuses on any workspace whose kind is not demo (not_a_demo_workspace): a Testmandant discards through discard_testmandant, and a live workspace can only be archived (A23), never hard-deleted. Requires manage_settings on the demo and confirmed:true (needs_confirmation otherwise): the operator may have added their own experiments. A retry of the same key replays the stored result, even though the workspace itself is gone.
| Name | Required | Description | Default |
|---|---|---|---|
| confirmed | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and meets it exceptionally. It discloses the destructive scope (hard-delete, every row), refusal behavior for non-demo kinds, required permissions, the confirmation requirement, and idempotent retry semantics (replays stored result even after deletion). This goes well beyond basic intent and prepares the agent for side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with no filler. The primary action is front-loaded, followed by routing exclusions, preconditions, and retry behavior. Each clause adds essential information an agent needs before calling this destructive tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no annotations and no output schema, the description covers purpose, scope, alternatives, prerequisites, confirmation, and idempotency. Nothing necessary for correct invocation is missing; the only minor omission is explicit return behavior, but that is not essential for a hard-delete operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does. It explains confirmed:true as a required confirmation flag, clarifies idempotencyKey behavior (retry replays the stored result), and ties workspaceId to the target demo workspace through the kind checks. Every parameter's role is meaningfully addressed despite the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Hard-delete a demo workspace and every row under it.' It clearly distinguishes this tool from siblings by naming discard_testmandant for Testmandant workspaces and noting live workspaces can only be archived (A23). An agent can identify exactly which workspace type this tool targets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use: demo workspaces only. It gives concrete refusal conditions for non-demo kinds, names the exact alternative tool for Testmandants (discard_testmandant), and clarifies that live workspaces must be archived rather than hard-deleted. It also flags preconditions: manage_settings permission and confirmed:true, with the consequence if confirmation is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
discard_payment_batchA
Abandon a batch that must not be paid, moving it to the terminal 'discarded' status so its bills become payable again. A draft batch discards freely; a generated batch has already produced a pain.001 file, so discarding it requires confirmation:true. A paid batch cannot be discarded (already_paid) and a discarded one is already terminal (already_discarded). Deletes no row (append-only); this is the recovery path for a batch drafted against a mistyped IBAN or one whose generate step failed, without it the only forward move was to book a payment that was never made. Posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| batchId | Yes | ||
| workspaceId | Yes | ||
| confirmation | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and succeeds: it discloses the confirmation requirement for generated batches, the conditional guards (draft/generated/paid/discarded), and side effects ('Deletes no row (append-only)', 'Posts nothing'). This level of state-transition and side-effect disclosure goes well beyond what annotations typically supply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded in the first clause, and the subsequent sentences are information-dense rather than padded. It is longer than strictly necessary, with clauses like the IBAN narrative adding context beyond the mechanical definition, but every sentence contributes to safe invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, and a state machine that depends on batch lifecycle, the description covers the purpose, all reachable states, the confirmation parameter, side effects, and concrete recovery scenarios. Nothing an agent needs in order to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% (no parameter is described in the schema), so the description must compensate; it does for the one non-obvious parameter by specifying 'confirmation:true' for generated batches. workspaceId, batchId, and idempotencyKey are left to their self-evident names and types, so the compensation is strong but not exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Abandon a batch that must not be paid'), the terminal outcome ('terminal 'discarded' status'), and the business effect ('its bills become payable again'), which cleanly separates it from payment-batch siblings like mark_batch_paid, generate_pain001, and payment_batch_transmit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use scenarios ('recovery path for a batch drafted against a mistyped IBAN or one whose generate step failed') and explicit when-not conditions embedded in the state machine ('A paid batch cannot be discarded (already_paid) and a discarded one is already terminal'). It contrasts with the forward path ('book a payment that was never made') but never names the alternative tool explicitly, which keeps it just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
discard_testmandantA
Discard a Testmandant and have it truly gone: a hard delete of the trial workspace and every row under it, because nothing in it is a real fiscal record until it goes productive. Refuses on any workspace whose kind is not sandbox (not_a_testmandant): a demo discards through G03, and a live workspace can never be hard-deleted through this family, so after go_productive this same call refuses. Once any step has committed to a live workspace, discard requires commit_migration. The plan itself survives and returns to planned with its trial-load state reset; the maps and declared controls are kept.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| confirmed | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description must reveal destructive behavior, refusal conditions, and side effects, and it does so thoroughly: hard delete, requires commit_migration after live commits, plan survives with reset trial-load state, maps and controls retained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the core destructive purpose, then adds conditional behavior and survival semantics. It is longer than minimal, but each sentence earns its place given the destructive and refusal logic.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of annotations and an output schema, the description covers the most important behavioral context: what is destroyed, what is refused, what survives, and what prerequisite applies. It is missing explicit per-parameter clarification and return/error shape, so it is not fully complete, but it is far above average for an operation of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for four undocumented parameters. It gives strong conceptual context about workspace kind and plan survival, but it never explicitly explains workspaceId, planId, confirmed, or idempotencyKey, so an agent still lacks key parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Discard a Testmandant' and clarifies this is a hard delete of the trial workspace and every row under it. It explicitly contrasts this with demo workspaces routed through G03 and with live workspaces that can never be hard-deleted through this family, so it is clearly distinguished from related sibling tools like discard_demo_workspace.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when the call refuses: non-sandbox workspace kinds, demo workspaces, live workspaces, and any workspace after go_productive. It also states when commit_migration becomes a prerequisite. This is strong 'when to use vs alternatives' guidance with concrete exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dispatch_previewARead-only
Render the concrete outbound message(s) as a pure read: nothing is sent, nothing is written. documentId (invoice/quote) yields exactly one resolved message; runId (dunning_run) yields one per debtor (narrowable via contactId), each with the recipient, locale, filled variables and attachment list; with neither, the saved or built-in text renders against MUSTER sample values (the editor's preview; locale picks the slot). A recipient without an email carries the owning send verb's own flag (needs_customer_email per A11, needs_email_transport per A15).
| Name | Required | Description | Default |
|---|---|---|---|
| runId | No | ||
| locale | No | ||
| contactId | No | ||
| documentId | No | ||
| workspaceId | Yes | ||
| documentKind | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses detailed behavioral nuances: exact message counts per mode, what fields are provided (recipient, locale, filled variables, attachment list), and the fallback to sample values. It also covers the special case of recipients without email and their associated flags, adding substantial context that the annotation alone could not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded with the most important behavioral caveat ('pure read'). All three sentences contribute information, using semicolons to pack multiple facts into a structured flow. The heavy use of domain jargon (MUSTER, A11, A15) may hinder parseability for agents, but the structure itself is logical and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers all three usage modes, specifies the output contents, and addresses an important edge case (missing email). Given there is no output schema, it provides reasonable detail for an agent to know what will happen. Some terms like 'MUSTER sample values' assume domain familiarity, and the return shape is described but not formally structured.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description must explain parameters, and it does explain documentId, runId, contactId, and locale with specific meaning. However, the two required parameters, workspaceId and documentKind, are not described at all; they are left entirely to the agent's domain knowledge. The description partially compensates for the schema gap but not completely.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Render') and resource ('outbound message(s)'), and immediately clarifies the key scope with 'nothing is sent, nothing is written'. It clearly differentiates the three input modes (documentId, runId, neither) and implies this is the preview counterpart to send operations, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong context for when to use the tool: to preview concrete outbound messages without sending. It explains the three different invocation patterns based on ID presence, effectively telling the agent which arguments select which behavior. However, it does not explicitly name alternative tools or state 'use this instead of X', leaving exclusions implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dispatch_text_upsertA
Speichere den Textbaustein (subject + body) for one document kind (invoice, quote or dunning_run) and locale: the workspace's own outbound message text, with {{variable}} interpolation validated against the kind's registered variable set AT SAVE TIME (an unknown variable is refused with unknown_variable naming the valid set, never a broken mail later). Naturally idempotent: it asserts the absolute state of one (documentKind, locale) slot, and a value-identical re-save writes nothing at all. Sends NOTHING and changes no send behaviour; a missing slot falls back to the built-in default text at preview.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | ||
| locale | Yes | ||
| subject | Yes | ||
| workspaceId | Yes | ||
| documentKind | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so richly. It discloses that {{variable}} interpolation is validated at save time, that unknown variables are refused with an unknown_variable error naming the valid set, that the operation is idempotent (value-identical re-save writes nothing), that it sends nothing and changes no send behavior, and that a missing slot falls back to built-in default text at preview. This is exemplary transparency for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: it front-loads the core purpose, then explains validation behavior, idempotency, and side-effect-free nature. Every sentence adds meaningful information. It is slightly long, but the length is justified by the behavioral complexity it discloses. The German opening ('Speichere') is a minor style inconsistency with the rest of the English description, but it does not harm clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (validation, idempotency, fallback behavior) and the absence of annotations and output schema, the description is remarkably complete. It covers what the tool does, what it validates, what it refuses, what it does not do, and what happens when a slot is missing. An agent has everything needed to call it correctly and predict its side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the meaning of documentKind (invoice, quote, or dunning_run), locale (the locale of the text), and subject/body (the text block being saved). It also explains the relationship between workspaceId and the slot (workspace's own outbound message text). It does not detail the exact format of locale values or the valid variable set, but the description adds substantial meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Speichere' = save/upsert) and a precise resource: the workspace's outbound message text (subject + body) for one document kind and locale. It explicitly names the allowed document kinds (invoice, quote, dunning_run), which distinguishes it from generic text or template tools. The scope is unambiguous and the description differentiates it from siblings like dispatch_preview and list_dispatches.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when to use this tool: to persist the workspace's own outbound message text for a specific document kind and locale. It also provides exclusions: it sends nothing and changes no send behavior, and a missing slot falls back to the built-in default text at preview. However, it does not explicitly name alternative tools for related operations (e.g., dispatch_preview for previewing, list_dispatches for listing), so the when-not-to-use guidance is implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
draft_generateA
Erstelle einen Antwortentwurf für einen Mail-Thread, lokal und in den Entwurfsordner des eigenen Mail-Programms (nie versendet: TILL hat kein Sende-Verb und kein SMTP). Der Entwurf spricht mit dem gelernten Schreibstil (E05) und wird NUR für einen Kontakt mit eingeschalteter Einwilligung (contact.ledger_grounding_enabled, via contacts_update) mit den Buchhaltungs-Fakten aus A16/A11/B00 fundiert; groundInLedger:false schaltet die Fundierung zusätzlich AUS, einschalten kann kein Parameter (die Einwilligungs-Asymmetrie). Ein unbekannter Absender fundiert nie; ohne read_sales fundiert nichts. Antwortet nothing_to_reply_to wenn die neuste Nachricht schon von Ihnen ist, needs_voice_profile ohne Schreibstil, needs_local_runtime ohne lokale Engine (kein Cloud-Fallback existiert), source_changed wenn die Nachricht sich geändert hat. Ein Schlüssel schreibt genau EINEN Entwurf.
| Name | Required | Description | Default |
|---|---|---|---|
| threadId | Yes | ||
| profileId | No | ||
| workspaceId | Yes | ||
| groundInLedger | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly states the tool never sends (no SMTP), writes to the drafts folder, grounds only with consent, uses idempotency keys, and lists all error conditions. This is thorough and transparent about side effects and failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and long, packed with many clauses and conditions in a single paragraph. While information-rich, it lacks structure (e.g., bullets) and could be more scannable. It is not excessively verbose, but it could be better organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers many edge cases: consent requirements, grounding conditions, error responses, and idempotency. It does not explicitly mention the success return value (e.g., draft ID), but given the complexity of the tool, the coverage is quite thorough. A small gap remains regarding what is returned on success.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters. It clarifies groundInLedger (turns grounding off when false) and idempotencyKey (one key = one draft), but does not explain workspaceId, threadId, or profileId. These are likely self-explanatory, but the description only partially compensates for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool creates a reply draft for an email thread, locally, and never sends it. It specifies the resource (mail thread) and the action (create draft), and distinguishes itself from siblings like draft_regenerate by focusing on creation. The verb and resource are specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear conditions for use (consent enabled, unknown sender, read_sales, etc.) and enumerates error responses. However, it does not explicitly name alternative tools or state when to prefer another tool (e.g., mail_draft_write). The context is sufficient to guide an agent, but explicit exclusions are missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
draft_listARead-only
Die Entwurfs-Läufe (P5), neuste zuerst, je Thread oder über den Arbeitsbereich: welches Modell welchen Entwurf schrieb (runtime, model_ref, prompt_sha256, nie der Prompt selbst), ob die Buchhaltung einbezogen war (grounded), der Status (ok, needs_local_runtime, needs_mailstore, failed), und der Entwurfstext ON DEMAND aus dem Entwurfsordner gelesen, nie aus SQLite. draftGone sagt, dass der Mensch ihn gesendet oder gelöscht hat; modelChanged, dass ein anderes Modell registriert ist als das, das ihn schrieb.
| Name | Required | Description | Default |
|---|---|---|---|
| threadId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=trueholiday. The description adds substantial behavioral context beyond that: the draft text is read on demand from the drafts folder 'nie aus SQLite', the prompt itself is never returned, and the semantics of the draftGone and modelChanged flags are explained. This gives the agent a precise mental model of what the tool will and will not do.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense run-on sentence stuffed with parenthetical enumerations (fields, statuses, flag meanings). Every clause carries information, so there is no waste, but the structure is hard to parse and front-loads only partially; the core purpose appears first, yet the flag semantics trail in a long tail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully explains the meaning of returned fields, statuses, and edge-case flags—the hardest part of invoking this tool. However, it never explicitly states that workspaceId is required or documents the expected parameters in a structured way, leaving an agent to reverse-engineer the call signature from prose.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It indirectly explains the workspaceId vs threadId scoping ('je Thread oder über den Arbeitsbereich') but never names either parameter explicitly, nor documents requiredness or formats. Partial compensation at best for the 2 undocumented parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb plus resource: it lists draft runs ('Entwurfs-Läufe'), sorted newest first, scoped per thread or workspace, and enumerates the fields returned. It is clearly a listing operation distinct from siblings like draft_generate or draft_regenerate, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies two usage scopes ('je Thread oder über den Arbeitsbereich'), which maps to the threadId/workspaceId parameters, but provides no guidance on when to prefer this tool over alternatives, no exclusions, and no prerequisites. All usage context must be inferred from the prose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
draft_regenerateA
Erstelle einen Entwurf neu, mit optionalem Hinweis ("kürzer", "förmlicher"): schreibt eine NEUE draft_run-Zeile (die Versuchsfolge bleibt nachvollziehbar) und ERSETZT die Nachricht im Entwurfsordner statt eine zweite daneben zu legen. Ein Entwurf, den der Mensch schon gesendet oder gelöscht hat, antwortet draft_gone und wird nicht neu erzeugt: gesendet heisst seiner. Einwilligung und RBAC werden neu gelesen, nie vom früheren Lauf geerbt.
| Name | Required | Description | Default |
|---|---|---|---|
| hint | No | ||
| draftRunId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses that a new draft_run row is created, the message is replaced, consent and RBAC are re-read (not inherited), and the draft_gone error for sent/deleted drafts. This is thorough and adds value beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured, front-loading the action and then detailing key behaviors. Each sentence adds information, though it could be slightly more concise. The use of capitalization for emphasis aids readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers core behavior, replacement semantics, failure conditions, and consent/RBAC re-reading, but lacks explanation of required parameters and does not describe the success response. Given no output schema, the absence of a success return description is a notable gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters. It only mentions the optional 'hint' with example values, but does not explain workspaceId, draftRunId, or idempotencyKey. The required parameters remain undocumented, leaving the agent to infer their meaning from names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it regenerates a draft ('Erstelle einen Entwurf neu'), with an optional hint, and explains the replacement behavior. It distinguishes itself from simply creating a second draft, but does not explicitly name sibling tools like draft_generate, so differentiation is implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides usage conditions: a draft already sent or deleted returns draft_gone and is not regenerated, and it replaces the message in the drafts folder. This gives partial guidance on when to use, but does not explicitly contrast with alternatives or state when to prefer this over draft_generate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ebill_delivery_statusARead-only
Track eBill deliveries: one by deliveryId (with its mirrored partner-status event trail), or a filtered list by invoiceId, local status (prepared|submitting|transmitted|failed), mirrored partner status (NWP_PENDING|OPEN|APPROVED|REJECTED|COMPLETED), or created-at range, newest first. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| status | No | ||
| invoiceId | No | ||
| deliveryId | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| partnerStatus | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds 'Read-only' redundantly but consistently. It adds value by disclosing 'newest first' ordering, the ability to fetch a single delivery with a 'mirrored partner-status event trail', and lists the exact status enums. These details go beyond the annotation and help predict behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that front-loads the core purpose, then enumerates the modes and filters. It uses semicolons to separate clauses, and every phrase adds information. There is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 8 parameters and no output schema. The description explains the main filters and ordering but does not describe the response format or the exact meaning of the 'mirrored partner-status event trail'. It also omits pagination details and does not mention that workspaceId is required. For a read-only tracker it is reasonably complete but leaves notable gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains deliveryId, invoiceId, local status (with enums), partner status (with enums), and created-at range (implied from/to), but does not explicitly map these to the schema property names (e.g., 'from'/'to'). It also omits savedViewId and does not mention the required workspaceId. It covers the main filters but leaves some parameters unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Track') and resource ('eBill deliveries'), and clarifies two modes: lookup by deliveryId or filtered list by various criteria. It includes specific enum values and ordering ('newest first'), and clearly distinguishes from generic 'delivery_status' by the 'eBill' prefix. An agent can immediately understand the tool's scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for eBill delivery tracking, but does not explicitly mention alternatives or when not to use it. It does not reference sibling tools like 'delivery_status' or 'ebill_transmit' to guide selection. The context is clear but exclusions are absent, so only implied usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ebill_prepareA
Turn an issued invoice (issued, sent, or partially_paid) into an eBill delivery payload on disk, carrying A11's QR reference unchanged. Files the A11 PDF as an E00 document (entity_kind ebill_delivery, OR 958f retention) and records an ebill_deliveries row (status prepared) with the payload's conformance facts (PDF/A profile, eBill addressing, byte length) recorded, never claimed. Refuses a draft/cancelled/settled invoice (invalid_state) and an invoice with no payment part (needs_qr_bill), writing nothing. Idempotent by outcome: at most one active (non-failed) delivery per invoice, so a re-prepare returns the existing row and its stored artifact regardless of key. Posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| invoiceId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: it discloses side effects (file creation, DB row), that conformance facts are 'recorded, never claimed', idempotency by outcome, and that refusal paths write nothing. This is notably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the main action, and every sentence carries information. However, it uses heavy domain jargon (A11, E00, OR 958f, conformance facts) and long compound sentences, which slightly reduce readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description covers input eligibility, failure modes, side effects, and idempotency. It partially describes the success return ('returns the existing row and its stored artifact'), which is enough for invocation, though the exact success response format is not stated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds real meaning to invoiceId (must reference an eligible invoice) and idempotencyKey (outcome-based idempotency, existing row returned 'regardless of key'). workspaceId is left to context, but it is self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it turns an issued invoice into an eBill delivery payload on disk, files the PDF as an E00 document, and records an ebill_deliveries row. It also distinguishes itself from the sibling ebill_transmit with 'Posts nothing.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit eligible invoice states (issued, sent, partially_paid) and refusal conditions (draft/cancelled/settled, no payment part), plus idempotency behavior. It does not explicitly name ebill_transmit as the alternative for actual sending, but 'Posts nothing' implies this is the prepare step.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ebill_transmitA
Transmit a prepared delivery through the owner-gated cloud connector (P8: outbound, draft-by-default; pass confirmed=true or enable the workspace dial). Guard order: no connector returns the honest {transmitted:false, reason:'cloud_tier'} and the artifact stays downloadable; then needs_biller_pid; then needs_confirmation; then payload_not_conformant if the payload is not PDF/A-3b, not eBill-addressed, or over 10 MB (nothing non-conformant is ever transmitted). On acknowledgement records the business case and drives the invoice issued->sent (A10) only from issued. Idempotent: a transmitted delivery never resubmits; transmission is at-least-once via a stored correlation id. Posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| confirmed | No | ||
| deliveryId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It thoroughly explains the tool's behavior: outbound, draft-by-default, idempotent (never resubmits, at-least-once via stored correlation id), drives the invoice state from issued to sent only from issued, records the business case on acknowledgement, and posts nothing. It also discloses failure modes and the exact reason codes, which is exceptional for agent decision-making.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and detailed, with every sentence contributing critical behavioral or precondition information. It is front-loaded with the primary action and then lists guard order and side effects. While it is longer than average, the structure is logical and the information is necessary given the tool's complexity and lack of annotations. It could be slightly more concise by removing jargon like 'P8' and 'A10', but it remains well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (multiple preconditions, failure modes, idempotency, state changes) and the absence of annotations or an output schema, the description provides remarkably complete context. It covers prerequisites, failure reasons, idempotency semantics, side effects, and what it does not do (posts nothing). An agent has sufficient information to call it correctly and understand the consequences. The only minor gap is the success return value, but that is a small omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It adds meaning to 'confirmed' (pass confirmed=true) and implicitly to 'idempotencyKey' (via stored correlation id for at-least-once). However, it does not explicitly explain 'workspaceId' and 'deliveryId' beyond the obvious context of transmitting a specific delivery in a workspace. While it provides some semantic value, it is not comprehensive for all four parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Transmit a prepared delivery through the owner-gated cloud connector'. It specifies the resource (a prepared delivery), the mechanism (owner-gated cloud connector), and implicitly the eBill context via the guard conditions and sibling tools like ebill_prepare. It distinguishes itself from ebill_prepare (prepares) and ebill_delivery_status (checks status) by focusing on the transmission step and its side effects.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage conditions: 'pass confirmed=true or enable the workspace dial' and a detailed guard order listing preconditions that must be met (no connector, needs_biller_pid, needs_confirmation, payload_not_conformant). It does not name alternative tools explicitly, but the guard order effectively tells the agent when it will reject the call. It lacks an explicit 'do not use when' statement, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
egress_self_testBRead-only
Run the offline proof (OP6): install a socket-level probe over every network vector (TCP, UDP, DNS, TLS, HTTP(S), fetch, WebSocket, subprocess) and run a REAL end-to-end draft generation (E04 mail read -> E05 voice retrieve -> E06 compose -> E04 Drafts write) under it. Reports honestly: passed:true with socketsOpened:0 when the whole loop opened zero sockets (the claim, measured, and it holds with wifi on or off because the probe counts sockets and does not depend on the network being down); egress_violated with the offenders named when any socket was opened (we report our own violation loudly, never degrade the claim quietly); needs_setup naming what is missing (mail_store, voice_profile, local_runtime, draftable_thread) when a real run cannot be attempted; self_test_incomplete when the loop could not finish for a non-egress reason. Produces a draft in the practitioner's own local Drafts folder; mints nothing else. Reads egress.read; deliberately readable by an agent auditing the claim.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint=true is contradicted by the description's explicit statement that the tool 'Produces a draft in the practitioner's own local Drafts folder' – a clear write side effect. This is a direct annotation contradiction, so the description fails to align with structured metadata.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a dense, single block with multiple semicolon-separated clauses. It is front-loaded with the core action but becomes long and winding when explaining all possible outcomes and side effects. While every sentence carries information, the structure could be clearer and more concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description thoroughly explains the tool's behavior, return outcomes (passed, egress_violated, needs_setup, self_test_incomplete), and side effect of producing a draft. However, it omits any explanation of the workspaceId parameter, and the contradiction with readOnlyHint undermines overall completeness for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one required parameter workspaceId with zero description coverage (0%), and the description does not mention this parameter at all. Since schema coverage is low, the description was expected to compensate but provides no meaning beyond the raw schema type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Run the offline proof (OP6)' that probes network vectors and performs an end-to-end draft generation. It clearly distinguishes this from a simple status check like the sibling egress_status by describing a full self-test with concrete outcomes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for auditing the egress claim ('deliberately readable by an agent auditing the claim') and explains what scenarios lead to which report. However, it never names sibling tools (e.g., egress_status) or states when to prefer this over alternatives, so it lacks explicit exclusions or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
egress_statusARead-only
The standing trust indicator (P5, observed never stored): the egress state of THIS process. state=local when the probe is installed and has observed zero outbound sockets this session, violated when it has seen one or more (with the offenders), unknown when the probe could not install (an unverified claim renders as unverified, never as optimism). socketsOpened is the observed count and since is when the observer started. Scoped to the TILL process, NOT the Mac: the mail app uses the internet (that is how mail arrives); this reports only what TILL itself did. Reads egress.read.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses the observed-never-stored nature, the state semantics (including the unverified claim renders as unverified, never as optimism), the counters (socketsOpened, since), and the read source ('Reads egress.read'). This fully explains what the agent will observe and what will happen when the probe can't install.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and information-rich but every clause earns its place. The parenthetical meanings and the explicit NOT distinguish this from sibling status tools. It's a bit long, but each detail (P5, never stored, state meanings, read source) is necessary for correct interpretation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the readOnlyHint and this tool's simplicity (1 param, no output schema), the description fully covers the state space, the observability caveats, and the scope. An agent can decide whether to call this tool and correctly interpret its result without ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema only lists workspaceId with zero description. The description does not explicitly explain the workspaceId parameter, but it deeply explains the output/semantics of the tool. With only one parameter, schema coverage is 0%, so the description could have mentioned that workspaceId specifies which workspace's process this queries. However, since the tool is about THIS process, the workspaceId role remains slightly understated; hence 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states this is a trust indicator of the egress state of THIS process, with the exact meaning of each state (local, violated, unknown). It specifies the resource ('TILL process, NOT the Mac') and distinguishes from the mail app. It's specific and unambiguous about what it reports.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly scopes the tool: 'Scoped to the TILL process, NOT the Mac: the mail app uses the internet... this reports only what TILL itself did.' It clarifies when this tool is relevant (checking this process's egress state) and what it does NOT measure, preventing confusion with other status/monitoring tools like delivery_status, sync_stream_status, or runtime_status.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
enable_automation_ruleA
Let a rule fire again. Requires manage_automations, which is the asymmetry that makes disabling safe to leave open: anyone may stop a rule, only an administrator may start one. An archived rule refuses.
| Name | Required | Description | Default |
|---|---|---|---|
| ruleId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the permission requirement, the asymmetry that makes disabling safe, and the failure condition for archived rules. This is meaningful behavioral context, though it does not mention idempotency or return behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with the core action front-loaded and no wasted words. It efficiently conveys the primary purpose, a key permission constraint, and a failure case.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple enable operation, the description covers the essential behavioral aspects (permission, failure on archived) and is reasonably complete. The only notable gap is parameter explanation, which is minor given the self-descriptive names. No output schema reduces the need to describe returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, and the description provides no information about ruleId or workspaceId. The parameter names are self-explanatory, but the description adds no additional meaning, leaving a gap that should be filled given the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Let a rule fire again' clearly indicates the tool re-enables a previously disabled rule. It distinguishes from siblings like disable_automation_rule and archive_automation_rule by focusing on the enabling action, though it does not explicitly use the word 'enable'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It specifies the permission requirement (manage_automations) and explains the asymmetry with disabling, giving clear when-to-use guidance. It also states that archived rules refuse, which warns against using this tool on archived rules. However, it does not explicitly contrast with update_automation_rule or other alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
enable_pluginB
Aktiviere eine Erweiterung (US-G02.2/3): register every capability the still-persisted manifest declares (its MCP tools, Studio screen, report source, automation actions) and start its sandbox, flipping status to installed. Refuses an incompatible plugin (plugin_incompatible, no run-anyway override, §6b) and a capability whose name another plugin has since taken (capability_name_conflict).
| Name | Required | Description | Default |
|---|---|---|---|
| pluginId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the registration scope (MCP tools, Studio screen, report source, automation actions), sandbox start, status flip, and two refusal modes. It could add idempotency behavior or permission requirements, but for an unannotated tool this is substantial transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the action and contains useful detail, but internal references like 'US-G02.2/3' and '§6b' add noise without explanation. The density of jargon ('still-persisted manifest', 'capability') makes it less accessible than it could be, though no sentence is outright wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and 0% schema description coverage, the description covers core behavior and failure modes well. However, it omits idempotency semantics despite a required idempotencyKey parameter, and does not describe the success result or status representation. It is workable but has clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description adds almost no parameter meaning. It refers to 'plugin' but never ties pluginId, workspaceId, or idempotencyKey to their roles, and the required idempotencyKey is not mentioned at all. The agent cannot learn what these parameters mean from either the schema or the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Aktiviere eine Erweiterung') and details exactly what the tool does: register every capability from the still-persisted manifest, start its sandbox, and flip status to installed. It distinguishes itself from sibling plugin tools by referencing the persisted-manifest state and naming concrete refusal conditions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context is implied rather than explicit: enabling a plugin whose manifest persists, after an earlier install step. However, it does not name sibling tools such as install_plugin or disable_plugin, nor state clearly 'use this when X, not when Y'. The failure conditions help, but the usage guidance is not fully spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
end_recurring_scheduleA
End a schedule for good: ended is terminal and cannot be resumed (create a new schedule instead). Ending an already ended schedule settles to the same answer. Generated invoices are untouched.
| Name | Required | Description | Default |
|---|---|---|---|
| scheduleId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does more than the name alone: it discloses terminality, idempotence for already-ended schedules, the side-effect that generated invoices are untouched, and irreversibility. It could still add permission prerequisites or other side effects, but the key behavioral traits are clearly covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the core action, the terminal/idempotent behavior with an alternative, and the invoice side-effect. It is front-loaded and ruthlessly concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no annotations and no output schema, the description covers the most important contextual facts: irreversibility, idempotent re-invocation, and non-interference with invoices. It lacks authentication/permission guidance and returns, but given the tool’s low complexity, the picture is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needed to compensate, but it does not actually explain workspaceId or scheduleId beyond what their names already imply. The only hint is that 'schedule' refers to the schedule being ended, which adds minimal meaning over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb-resource pair: 'End a schedule for good', making the tool’s purpose unmistakable. It also distinguishes it from the sibling lifecycle tools pause_recurring_schedule, resume_recurring_schedule, and update_recurring_schedule by emphasizing terminality ('cannot be resumed').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly defines when to use the tool ('for good'), states the exclusion ('cannot be resumed'), and points to the fallback alternative ('create a new schedule instead'). An agent can route between this tool and pause/resume/recreate without ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
env_copyA
Copy data one-way from a source environment into a lower one, sanitized. scope is instance (every workspace) or mandate: (one workspace, others in the target untouched). sanitize is raw (verbatim), pseudonymize (deterministic PII masking, valid CH test IBANs, amounts intact unless scaleFactor is given) or structure_synthetic (structure kept, amounts and transactions replaced). Build-then-swap: a failed copy leaves the prior target intact. Refuses target=main and refuses a source whose tier rank is not strictly above the target (down-only). The SECRET FLOOR always neutralizes live access secrets (bank/EBICS/token) regardless of level; the owner-only retainSecrets override lifts it. Refuses a copy into the active env without force. Unconfirmed, returns the exact plan and changes nothing (P8).
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | ||
| scope | No | ||
| source | Yes | ||
| target | Yes | ||
| sanitize | No | ||
| confirmed | No | ||
| scaleFactor | No | ||
| workspaceId | Yes | ||
| retainSecrets | No | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full responsibility for behavior disclosure, and it does so thoroughly. It reveals build-then-swap atomicity, the SECRET FLOOR neutralization of live access secrets, the retainSecrets override, all refusal conditions, and the unconfirmed-plan behavior. This is exactly the kind of behavioral detail an agent needs for a high-risk mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place: it front-loads the core purpose, then explains parameter semantics and safety guarantees in a natural order. For a complex, destructive 10-parameter tool, this density is justified and remains readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description covers the essential operational context: direction, scope, sanitization modes, atomicity, secret handling, refusals, and confirmation behavior. It does not describe the success return payload or idempotencyKey semantics, but those are minor relative to the safety-critical information that is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must add meaning, and it does for the most important parameters: scope, sanitize, scaleFactor, retainSecrets, and confirmed, plus an indirect explanation of force. However, idempotencyKey is not mentioned, and the workspaceId parameter is only implied through the 'mandate:<workspaceId>' scope, leaving some of the 10 parameters underspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: 'Copy data one-way from a source environment into a lower one, sanitized.' It also makes the direction and sanitization semantics explicit, which distinguishes env_copy from sibling tools like env_switch, env_reset, or env_create. The rest of the description reinforces the scope and safety model.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear preconditions and refusal rules: down-only tier direction, target=main refusal, active-env requires force, and unconfirmed calls return a plan. These tell an agent when the tool will or will not work. It does not explicitly name alternatives (e.g., 'use env_switch to change the active environment instead'), but the usage context is still strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
env_createA
Create a new environment and populate its data root per its policy: synthetic (seeded, seed selectable, default the Seeblick golden ledger) or live (an empty ledger). policy=copy is Phase B. Refuses a data root that collides with an existing environment or aliases main`s volume. Unconfirmed, returns a plan and changes nothing (P8).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| seed | No | ||
| scope | No | ||
| dbPath | No | ||
| policy | Yes | ||
| source | No | ||
| sanitize | No | ||
| tierRank | No | ||
| confirmed | No | ||
| guardTier | No | ||
| codeChannel | No | ||
| scaleFactor | No | ||
| workspaceId | Yes | ||
| runtimeTarget | No | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the burden of behavioral disclosure. It covers the plan-mode behavior (unconfirmed returns a plan, changes nothing), refusal on data-root collisions, and policy-dependent seeding. It does not explain post-confirmation behavior or output shape, but the core safety semantics are clearly disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences front-load the core purpose and behavior with little waste. However, internal jargon such as 'Seeblick golden ledger' and 'P8' is not expanded, which costs some clarity in an otherwise compact definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter tool with no output schema and no annotations, the description is not complete enough. It explains policy and plan mode but omits required parameter semantics, confirmation flow, return value, idempotency behavior, and the relationship to env_list/env_switch/env_reset in the sibling set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 15 parameters, so the description must compensate. It adds real meaning for policy, seed, and indirectly for confirmed and dbPath, but leaves most parameters (workspaceId, name, scope, source, sanitize, tierRank, guardTier, codeChannel, scaleFactor, runtimeTarget, idempotencyKey) unexplained. An agent would struggle to fill the rest of the schema correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Create a new environment') and distinguishes the two supported policies (synthetic vs live) with concrete effects on the data root. It also signals that policy=copy is not yet available, which prevents confusion with the sibling env_copy tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives actionable guidance on which policy to use and what each policy does, and notes that copy is Phase B. It does not explicitly name alternatives like env_copy or env_reset, but the Phase B remark effectively steers an agent away from unsupported copy behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
env_currentBRead-only
The active environment for this face, and whether it is read-only (a served or protected env is read-only from a local face).
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark readOnlyHint=true, and the description adds meaningful nuance by explaining that served/protected environments are read-only from a local face. It does not contradict the annotation and clarifies why read-only can be true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence with no filler, front-loaded with the main purpose. Slight jargon ('face', 'served/protected env') costs some clarity but does not make it bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a simple read-only getter: it states what is returned (active environment and read-only status) and the readOnlyHint covers the safety profile. It could name the exact return fields, but nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single required workspaceId has 0% schema description coverage and the description never mentions it. While workspaceId is self-explanatory, the description adds no semantic value for the parameter itself.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly identifies the resource (the active environment for a face) and the key returned facet (read-only status). It is more specific than env_list or env_status, though it does not explicitly contrast itself against those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use env_current over env_status, env_list, or the other env_* tools. The read-only explanation adds context but no selection guidance or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
env_deleteA
Remove a named environment`s data root and its landscape entry. Refused on main always, on a standard tier without force (use reset), and on the active environment (switch away first). Unconfirmed, returns a plan and changes nothing (P8).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| force | No | ||
| confirmed | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of disclosing behavior. It explains destructive scope, refusal conditions, the role of force, and that an unconfirmed call returns a plan and changes nothing. However, the cryptic 'P8' is unexplained, and it doesn't state permissions needed or what happens on a confirmed successful deletion.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences deliver the action, refusal conditions, and dry-run behavior with no filler. The primary purpose is front-loaded before the constraints and fallback guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema and minimal annotations, the description covers the core operation, critical refusals, the alternative tool, and unconfirmed behavior. It remains slightly incomplete because 'P8' is unexplained and idempotencyKey's role is absent, but overall an agent has enough to decide whether and how to call this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the five parameters. It clarifies the purpose of 'name', 'force', and 'confirmed' indirectly, but it does not explain 'workspaceId' or 'idempotencyKey' at all, and the semantics of 'confirmed' are only implied rather than stated explicitly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Remove') with a precise resource ('named environment's data root and its landscape entry'), making the operation unmistakable. It also implicitly differentiates from sibling tools like env_reset by describing a destructive delete rather than a reset.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when the tool is refused: on main always, on standard tier without force, and on the active environment. It even names the alternative ('use reset') and prescribes the required precondition ('switch away first'). This gives an agent clear routing and guardrail guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
env_listARead-only
List every TILL environment in the landscape with its code channel, data policy, guard tier, runtime target, tier rank, last refresh, size and which one is current. Standard tiers (main, test, develop) are pinned first, then named environments. Fails loud if the landscape control file fails its integrity check.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses ordering behavior ('Standard tiers ... pinned first, then named environments') and a failure mode ('Fails loud if the landscape control file fails its integrity check'). This adds useful behavioral context that the annotation alone does not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and then lists specific output fields, followed by ordering and failure behavior. Each sentence carries relevant operational detail, though the list of fields is slightly long; overall it is appropriately concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output fields, ordering, and failure behavior are well covered, and the readOnlyHint annotation handles safety. However, the required workspaceId parameter is left completely unexplained, creating ambiguity about whether the tool lists environments for a specific workspace or across the entire landscape. This is a notable completeness gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With one required parameter and 0% schema description coverage, the description carries the full burden for explaining workspaceId, but it never mentions the parameter, its role, format, or how it relates to the 'landscape'. The description adds no meaning to the input schema beyond its existing name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb, resource, and scope: 'List every TILL environment in the landscape'. It then enumerates the exact fields returned, which makes the tool's purpose concrete and distinguishable from siblings like env_current or env_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies its usage through 'List every TILL environment', but it does not explicitly state when to prefer this tool over env_current, env_status, or other environment-related siblings. No when-not guidance or alternative routing is provided, so usage is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
env_resetA
Wipe and repopulate an environment from its data policy, build-then-swap: a new db file is seeded and gated (foreign-key check plus a balance re-foot) and swapped in only on success, so a failed reset leaves the prior environment intact. Refused on main (by data root, not just name) and on the active env without force. Reset of a copy-policy env is Phase B. Unconfirmed, returns a plan and changes nothing (P8).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| seed | No | ||
| force | No | ||
| confirmed | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries the behavioral burden. It discloses the build-then-swap mechanism (atomic swap), safety (failed reset leaves prior environment intact), restrictions (main/active without force), and unconfirmed behavior (returns a plan, changes nothing). This is comprehensive and transparent for a destructive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the core action, then safety, restrictions, and unconfirmed behavior. It uses domain-specific jargon like 'build-then-swap' and 'balance re-foot' which may be cryptic but are efficient for the intended audience. Slightly cryptic terms like 'P8' reduce clarity but do not undermine structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers key safety and usage aspects but omits important context: it does not describe the return value (no output schema exists), nor does it explain several parameters (seed, idempotencyKey, workspaceId). It also does not specify what a confirmed reset returns or any prerequisites. Given the destructive complexity, more detail is needed for full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains 'force' (required to reset active env) and 'confirmed' (unconfirmed returns plan), but does not explain 'seed', 'workspaceId', 'idempotencyKey', or 'name' beyond their existence. Only two of six parameters are clarified, leaving significant gaps for an agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action: 'Wipe and repopulate an environment from its data policy.' This is a specific verb-resource combination that distinguishes it from siblings like env_delete (delete) or env_copy (copy). The purpose is unambiguous and immediately understandable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear constraints on when the tool can be used: it is refused on main and on the active environment without force, and unconfirmed calls return a plan without changes. It also notes that copy-policy resets are Phase B. However, it does not explicitly name alternative tools or state when to prefer env_reset over others, though the restrictions give practical usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
env_statusARead-only
Detail one environment: its data freshness, guard tier, runtime target, recorded code channel and drift against the built one, and whether its data root exists. Fails loud on a landscape integrity mismatch.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already signals safety, but the description adds meaningful behavioral context beyond that: it names what is inspected and explicitly warns that the tool 'Fails loud on a landscape integrity mismatch'. That failure-mode disclosure is valuable and not present in structured data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core operation, followed by a compact list of inspected attributes and a terminal failure-mode warning. Every sentence earns its place with no fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only two-parameter status tool with no output schema, the description covers the main contract: what data is returned and the notable failure behavior. It does not fully define the term 'landscape integrity mismatch' or describe the exact response shape, but the listed fields provide enough operational context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the description does not explicitly explain the workspaceId and name parameters or their formats/constraints. The phrase 'Detail one environment' weakly implies name identifies the environment, but with no schema descriptions the description should compensate more than it does.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Detail one environment') and enumerates the exact aspects returned: data freshness, guard tier, runtime target, code channel, drift, and data-root existence. This clearly distinguishes it from list/switch/current/create-style environment siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The tool's purpose is clear enough that an agent can infer when to use it: when details about a single named environment are needed. However, it never explicitly contrasts it with env_list, env_current, or env_switch, leaving the when-to-use-vs-alternatives decision implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
env_switchA
Set the active environment for this face. Switching to a served or protected env (main) records the pointer and marks the face read-only; destructive verbs stay refused on main regardless. Unconfirmed, returns a plan and changes nothing (P8).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| confirmed | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full responsibility and delivers: it discloses side effects (records pointer, marks face read-only), safety constraints (destructive verbs refused on main), and the non-destructive unconfirmed path. This is exemplary transparency for a state-changing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, no filler. The main action is front-loaded, and each subsequent sentence adds distinct behavioral detail without redundancy. Optimal length for the information conveyed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers key behavioral aspects (read-only marking, destructive refusal, confirmation behavior) but leaves gaps: what happens when confirmed (presumably executes), what the returned plan contains, the meaning of 'P8', and the roles of workspaceId and idempotencyKey. Without an output schema or annotations, these omissions create a moderate gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description only indirectly explains the 'confirmed' parameter via 'Unconfirmed, returns a plan and changes nothing'. It provides no meaning for 'workspaceId', 'name', or 'idempotencyKey'. The description fails to compensate for the complete lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb ('Set') and resource ('active environment for this face'), and adds distinguishing specifics: switching to a served/protected env ('main') records the pointer and marks read-only, while destructive verbs remain refused. This clearly separates it from sibling tools like env_create, env_copy, or env_reset.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly tells when to use this tool (when you need to change the active environment) and explains the confirmation flow ('Unconfirmed, returns a plan and changes nothing'), which is critical usage context. It does not explicitly compare to alternatives like env_create or env_reset, but the uniqueness of 'switch' makes the context clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
expense_claim_approveA
Approve a submitted claim: POSTS the reimbursement liability via A02 (Dr per-category expense + Dr 1170/1171 Vorsteuer; Cr 2000 Kreditoren gross, the A17 vendor-bill shape with the employee as counterparty). P8 draft-gated: WITHOUT confirm:true it returns the journal-entry preview and posts NOTHING. Requires spesen.approve AND post. Refuses self_approval (approving your own claim, four-eyes, including an automation firing), period_locked (a locked posting period, pre-checked, no partial post), and invalid_transition (not submitted). Re-approving never double-posts. needs_employee_contact if the employee has no vendor-role contact to settle against. CONSEQUENCE: Approves the expense claim and posts the reimbursement liability to the ledger.
| Name | Required | Description | Default |
|---|---|---|---|
| claimId | Yes | ||
| confirm | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden and does so exceptionally. It reveals the draft-gated behavior without confirm:true, the exact journal-entry posting lines, idempotency on re-approval, permission requirements, refusal modes, and a missing-contact condition. This goes well beyond a generic 'approves' statement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-rich, with the core purpose front-loaded. It contains some jargon and a redundant capitalized CONSEQUENCE line, but every sentence adds operational value, so the length is mostly justified for a high-stakes posting action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex approval-and-posting tool with no annotations and no output schema, the description covers the critical context: draft gating, permissions, refusal cases, idempotency, and contact prerequisites. It does not explicitly describe the success return value after confirm:true or define each parameter, which prevents a perfect score, but it is otherwise operationally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It adds meaning mainly to the confirm parameter and the idempotency concept, but it does not explicitly explain claimId, workspaceId, or idempotencyKey beyond indirect implications like 'Re-approving never double-posts.' This is partial compensation, not complete parameter-level guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Approve a submitted claim' and then specifies what that means: it posts the reimbursement liability to the ledger via A02 with the employee as counterparty. It clearly differentiates itself from related expense-claim siblings like submit, reject, and reimburse by naming the exact posting shape and workflow stage.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: approval is for submitted claims, requires 'spesen.approve AND post', and refuses invalid_transition, self-approval, and locked periods. It does not explicitly name alternative tools, but the workflow conditions and the confirm:true preview-vs-post distinction effectively tell the agent when and how to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
expense_claim_createA
Create a draft expense claim (Spesenabrechnung) for an employee. Requires spesen.submit. Add lines with expense_line_upsert, then submit with expense_claim_submit.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| currency | No | ||
| employeeId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It clearly states the tool creates only a DRAFT (not submitted), requires a permission, and names the downstream steps in the lifecycle. It doesn't disclose side effects like idempotency behavior or validation rules, but the draft-state disclosure is valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with high information density: purpose, permission, and workflow. No redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a simple 5-param create-draft tool with no output schema, the description covers the permission prerequisite and the subsequent workflow steps. It does not clarify what idempotencyKey means or whether currency has defaults, but the required workflow context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so no parameter is explained in schema. The description does not explain workspaceId, employeeId, title, currency, or idempotencyKey. However, the names are mostly self-explanatory and the description establishes the context (employee, draft claim). Baseline 3 with no compensation for undocumented params is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Create a draft expense claim (Spesenabrechnung) for an employee.' The German term clarifies the domain object, and the name is unambiguous against siblings like expense_line_upsert and expense_claim_submit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states a required permission (spesen.submit) and gives a clear workflow: create draft, add lines with expense_line_upsert, submit with expense_claim_submit. This tells an agent exactly when to use it and what to call next.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
expense_claim_getBRead-only
Read one claim with its lines, the posted journal entry (once approved) and the payment (once reimbursed). Self-scoped: a bare reader reading a colleague`s claim id gets not_found.
| Name | Required | Description | Default |
|---|---|---|---|
| claimId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the read-only nature is covered by structured data. The description adds value by disclosing that the returned data depends on claim state ('once approved', 'once reimbursed') and reveals the authorization behavior (not_found for colleague's claims). However, it doesn't mention what happens for a claim that is approved but not reimbursed, or any other edge cases, though with annotations present this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. Front-loads the read action and returned components, then the security caveat. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read tool with readOnlyHint and no output schema, the description covers the resource, the conditional contents, and the authorization quirk. It doesn't explain what a 'claim' is or how the response is shaped, but given the simple schema and existing sibling nomenclature, this is roughly adequate, with a minor gap on workspaceId semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are two string parameters (workspaceId, claimId). The description mentions 'claim' and 'colleague's claim id', which hints at claimId's meaning, but it does not clarify workspaceId at all or provide the parameter-format context the schema lacks. Given zero schema coverage, the description should compensate more; it only partially does.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb ('Read') with resource ('one claim') and clarifies scope ('with its lines', 'posted journal entry once approved', 'payment once reimbursed'). It distinguishes itself from sibling expense_claim_list (which presumably lists claims) and expense_claim_create by being the getter. However, it doesn't explicitly name a sibling alternative, so a small differentiation gap remains.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Self-scoped' sentence conveys an important access caveat (a bare reader gets not_found for a colleague's claim id), which implies usage context. But it doesn't explicitly say when to use this tool instead of expense_claim_list, nor does it state prerequisites like requiring the claim id. The guidance is present but implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
expense_claim_listARead-only
List expense claims, SELF-SCOPED: only hr.manage (full cross-employee) or spesen.approve (the submitted-claims approval queue) widens the read beyond the caller`s own claims. status and employeeId narrow within scope.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| employeeId | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true, and the description adds scope-limiting behavior (self-scoped vs. permission-widened) and the filtering effect of status and employeeId. This complements the annotation without contradiction, giving an agent a clear understanding of what the read operation will return under different contexts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, information-dense sentence with no wasted words. It front-loads the core action, then efficiently explains scope modifiers. The use of technical permission names is compact, though it may be slightly cryptic to an agent without prior exposure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple (list with filters) and read-only, but the description omits details on workspaceId and savedViewId, and does not mention pagination, sorting, or output structure. The scope and filtering guidance is useful, but these gaps mean an agent may still make wrong calls when needing to use the full parameter surface.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverageached and four parameters, the description only mentions status and employeeId as narrowing filters, leaving workspaceId (required) and savedViewId completely unexplained. The agent cannot reliably infer the meaning or allowed values of these parameters, so the description fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a clear verb-object pair 'List expense claims' and immediately distinguishes itself by specifying self-scoping behavior and the exact permissions (hr.manage, spesen.approve) that widen read scope. This makes it easy to differentiate from other expense_claim_* tools like expense_claim_get or expense_claim_approve.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives practical context: by default only the caller's own claims are visible, and certain permissions broaden access. It implicitly tells an agent when this tool is appropriate versus those requiring extra permissions, though it does not explicitly name alternative tools or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
expense_claim_reimburseA
Reimburse an approved claim: pays the employee via A14 recordPayment (outgoing supplier settlement clearing 2000 Kreditoren, Dr 2000 / Cr bank) and prepares the pain.001 as a LOCAL artifact. P8 draft-gated: WITHOUT confirm:true it returns the payment + artifact plan and pays NOTHING. Transmission of the pain.001 to a bank is cloud-tier: the result carries transmitted:false, reason:cloud_tier. Idempotent: a reimbursed claim is never paid twice. Requires spesen.approve AND pay AND post. bankAccountId names the paying Bankkonto. CONSEQUENCE: Pays the employee for an approved expense claim; a reimbursed claim is never paid twice.
| Name | Required | Description | Default |
|---|---|---|---|
| claimId | Yes | ||
| confirm | No | ||
| workspaceId | Yes | ||
| bankAccountId | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it is exceptionally explicit: it discloses the accounting posting, that without confirm:true it 'pays NOTHING', that the pain.001 is only a local artifact, that result carries transmitted:false, reason:cloud_tier, that the operation is idempotent, and the required permissions. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The main action is front-loaded and most sentences add unique information, but the final CONSEQUENCE sentence repeats both 'pays the employee' and 'a reimbursed claim is never paid twice' from earlier. The description is dense and could be tightened without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, this covers purpose, side effects, preconditions, permissions, idempotency, draft behavior, and transmission limitations. The main gaps are the exact response/plan shape beyond transmitted:false and whether bankAccountId is required when confirm:true is supplied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains confirm via the draft-gate behavior, bankAccountId as the paying Bankkonto, idempotencyKey indirectly through idempotency, and claimId through context. However, workspaceId is left generic and no format details or conditional requirements are given for bankAccountId.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description immediately names the action ('Reimburse an approved claim'), the resource being acted on, and the concrete outcome ('pays the employee'). It also distinguishes itself from approval/submission tools and from generic payment-generation tools by stating it creates the pain.001 locally and is draft-gated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly sets preconditions: the claim must be approved, confirm:true is required to actually pay, specific permissions are required, and bankAccountId identifies the paying Bankkonto. It also states that bank transmission is not performed here via transmitted:false/cloud_tier, though it does not explicitly name an alternative sibling for real transmission.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
expense_claim_rejectA
Reject a submitted claim with a reason (reason_required if empty). Terminal (submitted -> rejected). An APPROVED (posted) claim is not rejected: it is corrected by an A02 reversing entry plus a fresh claim (invalid_transition otherwise). Requires spesen.approve (approve and reject share the review right).
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| claimId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it delivers: terminal transition, invalid_transition for approved claims, and permission requirement. It does not discuss idempotency or return behavior, but the key safety and state behaviors are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences front-load the action and state transition, then cover the exception and permission. No filler or redundant restatement of schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter mutation with no output schema and no annotations, the definition is close to self-sufficient: it specifies the target state, the invalid transition, the remediation path, and the permission. Only the idempotencyKey semantics and exact response/error format are left unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the reason field's validation ('reason_required if empty') and indirectly defines claimId as a submitted claim, but workspaceId and idempotencyKey are left entirely to naming convention.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Reject'), a resource ('a submitted claim'), and the required reason. It distinguishes itself by declaring the terminal state transition (submitted -> rejected) and by explaining the approved-claim alternative, making it clearly different from expense_claim_approve and requisition_reject.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when it applies (submitted claims) and when it must not be used (approved/posted claims), giving the corrective alternative: an A02 reversing entry plus a fresh claim. Also names the required permission (spesen.approve) and notes approve/reject share the review right.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
expense_claim_submitA
Submit a draft claim for approval. Refuses empty_claim (no lines) and receipt_required (a line over CHF 50 with no receipt, naming the offending line_ids). Snapshots the base total. Requires spesen.submit. Idempotent on the key.
| Name | Required | Description | Default |
|---|---|---|---|
| claimId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden of behavioral disclosure and meets it thoroughly. It discloses validation failures (empty_claim, receipt_required), names the offending line_ids in errors, mentions a side effect (snapshots base total), states the required permission (spesen.submit), and guarantees idempotency on the key. This is exceptional transparency for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured, front-loading the primary action, then detailing error conditions, side effects, permissions, and idempotency in separate sentences. Every sentence contributes new information without redundancy, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description covers the essential behavioral aspects an agent needs: what the tool does, when it refuses, what it requires, and its retry safety. It does not explain success return values or how to obtain a claimId, but those are minor given the tool's straightforward submit action and the presence of sibling list/get tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all parameter meaning. While it indirectly explains idempotencyKey via 'idempotent on the key,' it provides no meaning for claimId (beyond 'draft claim' in the purpose) or workspaceId. The description does not clarify formats, sources, or dependencies for two of three required parameters, leaving a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb-resource pair, "Submit a draft claim for approval," which precisely identifies the action and the object. It distinguishes itself from sibling tools like expense_claim_approve, expense_claim_reject, and expense_claim_reimburse by specifying the draft state and the approval goal, leaving no ambiguity about its role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool (for draft claims) and contrasts with later-stage siblings, but it does not explicitly state when not to use it (e.g., for already-submitted claims). It gives clear context through the terms 'draft' and 'for approval,' which is sufficient for an agent to select it correctly, though explicit exclusions are missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
expense_line_upsertA
Add or edit a claim line (only while the claim is draft). line carries expense_date, category (travel|meals|supplies|it|other), amountMinor (integer Rappen, the receipt total), currency (default CHF), an optional fxRate for a foreign line, an optional taxCode (an input-side Vorsteuer code, or none), an optional receiptDocumentId (an E00 file), and optional costCenterId/projectId/expenseAccountId. The per-line tax_code + base + tax are resolved ONCE by A05 and stored as values (MWSTG Art. 28). Requires spesen.submit. Pass line.lineId to edit.
| Name | Required | Description | Default |
|---|---|---|---|
| line | Yes | ||
| claimId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it delivers: it discloses the mutation, the draft-state precondition, the required permission, and the key persistence behavior that per-line tax_code + base + tax are resolved ONCE by A05 and stored as values (MWSTG Art. 28). Idempotency behavior is not explained, but the operational traits an agent needs are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense paragraph that front-loads the purpose and then proceeds logically through fields, tax behavior, permission, and the edit mechanism. Every sentence carries operational information; there is no filler or redundant restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with a nested object, 0% schema coverage, no annotations, and no output schema, the description covers almost everything needed to call it correctly: purpose, precondition, field semantics, permission, and edit behavior. The remaining gaps are the return value (what the upsert returns) and the precise semantics of idempotencyKey.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, yet the description fully documents the nested line object: expense_date, category with its enum (travel|meals|supplies|it|other), amountMinor as integer Rappen, currency default CHF, and each optional field (fxRate, taxCode, receiptDocumentId, costCenterId/projectId/expenseAccountId). Only idempotencyKey semantics are left to inference; the other top-level params (workspaceId, claimId) are self-explanatory by name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair ('Add or edit') applied to a specific resource ('claim line'), with the draft-state scope attached immediately. This cleanly separates it by resource granularity from the claim-level siblings such as expense_claim_create, expense_claim_submit, and expense_claim_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit preconditions and usage conditions: 'only while the claim is draft', 'Requires spesen.submit', and 'Pass line.lineId to edit' distinguishes the add case from the edit case. It does not name alternative tools or a when-not-to-use scenario, but the context is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_journalARead-only
The Buchungsjournal of a period ('YYYY-MM' or 'YYYY') as a locale-neutral CSV file, base64-encoded: one row per posted line exactly as the ledger holds it (integer Rappen in line and base currency, ISO dates, stored fx_rate and VAT trace values verbatim, OR Art. 958f: a faithful copy, never a recomputation), plus a footing total record. Same period, same bytes. Nothing is transmitted anywhere.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | ||
| period | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already signals safety, so the description earns credit by adding concrete behavior: exact row semantics ('one row per posted line'), fidelity guarantees ('faithful copy, never a recomputation'), determinism ('Same period, same bytes'), and the privacy guarantee ('Nothing is transmitted anywhere'). This goes well beyond the boolean annotation, though it does not mention error cases or size limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence with no filler: every clause carries meaningful detail (format, encoding, row content, fidelity, determinism, privacy). It is front-loaded with the core purpose. The phrasing is slightly run-on, especially the 'OR Art. 958f' aside, but it remains compact and informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description thoroughly explains what the caller receives: a base64 CSV, row structure, field values, and a footing record. It also covers determinism and data residency. Missing are details about the optional 'format' parameter and error handling for invalid periods, but for a read-only export with a rich description, it is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. It does add meaning for 'period' by specifying the two accepted formats ('YYYY-MM' or 'YYYY'), which is valuable. However, it fails to explain the 'format' parameter at all, and 'workspaceId' is left entirely to its name. This is a clear gap for the ambiguous optional parameter, warranting only a mid score despite good coverage of the period.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('export') and a specific resource ('the Buchungsjournal of a period') and fully specifies the output ('locale-neutral CSV file, base64-encoded'). It also names the period formats ('YYYY-MM' or 'YYYY') and differentiates from sibling export tools like export_statements, export_vat, and export_workspace by its focus on the journal/ledger. An agent can confidently know what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies the tool is for exporting the journal of a period, so an agent can infer when it would be appropriate. However, it gives no explicit guidance on when to prefer this tool over siblings like export_statements or export_vat, and no exclusions or alternative conditions are named. Usage guidance remains implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_statementARead-only
Render any of the four statements as a LOCAL file and hand back its bytes, base64-encoded: kind one of 'trial', 'balance', 'income', 'ledger', format one of 'csv', 'pdf', plus that statement's own parameters. The CSV is locale-neutral for re-import (integer Rappen, ISO dates, a record_type column carrying the totals) and the PDF is a print artifact with no PDF/A conformance claimed. The model is computed by the same verb the screen calls, so the file and the screen cannot disagree. Nothing is transmitted: e-filing or publishing the artifact is a cloud-tier concern and is not part of this tool.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| kind | Yes | ||
| format | Yes | ||
| groupBy | No | ||
| accountId | No | ||
| compareTo | No | ||
| periodEnd | No | ||
| periodStart | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond the readOnlyHint annotation by detailing behavioral specifics: the output is base64-encoded bytes, CSV is locale-neutral with integer Rappen and ISO dates, PDF makes no PDF/A conformance claim, and the model is computed by the same verb as the screen ensuring consistency. These details help the agent understand the exact nature of the output and its guarantees, adding value beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured, front-loading the core action and then providing concise clarifications about format details and behavioral guarantees. Each sentence adds distinct information without redundancy. While it is somewhat lengthy, it remains focused and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 9 parameters, a nested object, no output schema, and no schema descriptions, the description covers the main purpose and format semantics but omits explanations for several parameters and the expected return structure beyond base64 bytes. It also does not specify how parameters like periodStart and periodEnd are used or whether compareTo is required for certain statement types. This leaves gaps for an agent to correctly call the tool in all scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate for parameter meaning. It does enumerate valid values for 'kind' and 'format' and mentions 'plus that statement's own parameters,' implying the remaining parameters are context-specific. However, it does not explain workspaceId, asOf, groupBy, accountId, compareTo, periodEnd, or periodStart, leaving the agent to infer their roles. This is partial compensation at best.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action: 'Render any of the four statements as a LOCAL file and hand back its bytes, base64-encoded.' It specifies the resource (four statement kinds) and the output format, making the purpose unambiguous. It also distinguishes itself from related tools by emphasizing local-only generation and no transmission, which differentiates it from export_vat or export_statements.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool (for local file generation) and explicitly states what it does not do ('Nothing is transmitted: e-filing or publishing the artifact is a cloud-tier concern and is not part of this tool'), which guides agents away from using it for transmission. However, it does not explicitly name alternative tools or provide a clear when-to-use versus when-not-to-use scenario, leaving some inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_statementsARead-only
The filing pair for a period: the Bilanz as of the period end AND the Erfolgsrechnung over it, one artifact each, format 'pdf' or 'csv'. Each statement is rendered by A08's own export_statement engine (this tool only packages the pair), so the figures cannot differ from the screen or from export_statement, and A08's statutory caveats apply unchanged: do not present the Bilanz as the OR minimum structure.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | ||
| period | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds meaningful context: it does not render statements itself, uses A08's export_statement engine, guarantees consistency with the screen and export_statement, and carries statutory caveats. This goes beyond simply restating the name or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences front-load the core artifact definition and then add essential behavioral and statutory caveats. Every clause contributes something useful; there is no filler or padded explanation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the artifact composition, format options, rendering engine, consistency guarantee, and statutory caveat, all without an output schema. It does not specify the exact period string format or the shape of the returned artifacts, but for a read-only packaging tool this is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the burden. It meaningfully clarifies period ('as of the period end' and 'over it') and format ('pdf' or 'csv'), but does not explain workspaceId at all. That leaves one of three parameters undocumented despite the heavy burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool's purpose: it packages the statutory filing pair of Bilanz and Erfolgsrechnung for a period, one artifact each. It explicitly contrasts this with export_statement, noting this tool 'only packages the pair,' which distinguishes it from the sibling that renders individual statements.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'filing pair' framing tells the agent when this tool is appropriate, and the contrast with export_statement implies the alternative for single-statement use. It does not explicitly state when not to use related tools like export_journal or export_vat, but the package pair focus is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_vatARead-only
The MWST-Abrechnung figures of a period as a per-form-line CSV working paper: every Ziffer line exactly as vat_return computes it (stored trace values, never recomputed), plus the payable (Ziffer 500) and credit (Ziffer 510) totals. The file a Treuhänder lays beside the annual accounts; the ESTV ePortal upload artifact itself is vat_export_ech0217, not this.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | ||
| period | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses a non-obvious behavior: values are stored trace values from vat_return, never recomputed. It also clarifies output contents (every Ziffer line plus Ziffer 500/510 totals). No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences with no filler; the main output and scope are front-loaded, and the sibling distinction is placed at the end. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is rich for a read-only export: it explains output format, content lineage, and the correct alternative. Minor gaps remain around period/format value syntax, but the selection and invocation context is mostly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the three parameters. It only mentions 'of a period' and CSV output; it does not explain period format, workspaceId meaning, or the optional format property.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific deliverable (per-form-line MWST-Abrechnung CSV working paper) and explicitly contrasts it with vat_export_ech0217. An agent can tell exactly what this tool produces and what it is not.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the intended use ('file a Treuhänder lays beside the annual accounts') and explicitly names the alternative for ESTV ePortal upload (vat_export_ech0217). This gives both a when-to-use and a when-not-to-use rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_workspaceA
Export the whole workspace to a documented, open .tillexport bundle (one JSONL file per table plus a manifest and FORMAT.md) a human can read without TILL. Optional scope names a subset of tables. Not a restore source.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. It adds useful disclosure: output is a documented open bundle readable without TILL, optional scope narrows tables, and it is not a restore source. It could add idempotency or size/volume warnings for a whole-workspace export, but it conveys the essential behavior well beyond what structured fields show.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with zero fluff. The core behavior and format are front-loaded, the optional scope is in sentence two, and the crucial warning is the final sentence. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an export tool with no annotations and no output schema, the description covers input selection, output format, and a key exclusion. It does not mention idempotencyKey semantics or potential size/performance effects of exporting an entire workspace, which would help an agent set expectations, but overall it is solidly usable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and there are 3 parameters, so the description must compensate. It explains 'scope' well ('Optional scope names a subset of tables'). However, workspaceId and idempotencyKey receive no semantic treatment beyond their raw names in the schema; the description adds no guidance on what idempotencyKey is for or what workspace scope entails.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Export'), names the exact resource ('the whole workspace'), and specifies the output format ('.tillexport bundle' with one JSONL file per table plus manifest and FORMAT.md). It also distinguishes itself from a restore operation by explicitly stating 'Not a restore source.' In a large sibling list containing export_journal, export_statements, export_vat, and export_statement, this description makes it clear that it exports the entire workspace rather than a single report.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what the tool does but never says when to choose it over alternatives like export_journal, export_statements, or export_vat. The 'Not a restore source' line warns against one misuse but provides no explicit when-to-use or when-not-to-use guidance, leaving the agent to infer usage from the generic export framing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_deleteA
Erase a stored file and its bytes for good, once nothing forbids it: a retention date still in the future refuses with retention_locked (on a superseded version too, and against the recomputed statutory floor rather than only the stored date), and a version something supersedes refuses with not_head_version. Irreversible, so an agent caller gets {staged:true} and a pending badge unless it passes confirmed=true. Re-send the SAME idempotencyKey with confirmed:true to complete a staged deletion: confirmed is part of the replay identity, so the confirm is not swallowed as a repeat of the staging.
| Name | Required | Description | Default |
|---|---|---|---|
| fileId | Yes | ||
| confirmed | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses irreversibility, staged deletion, the pending badge, idempotency replay identity, the confirmed flag, and specific error conditions. This is far beyond minimal behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three dense sentences, each adding unique information: the operation and guards, irreversibility and staging, and the idempotent confirmation flow. There is no repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no output schema and no annotations, the description covers the success flow, staging, confirmation, failure modes, and idempotency. An agent has enough information to invoke it correctly and understand the consequences.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides zero parameter descriptions, so the description must compensate. It thoroughly explains confirmed and idempotencyKey, including the subtle point that confirmed is part of replay identity. fileId and workspaceId remain inferable from their names and context, but the critical safety parameters are well described.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Erase a stored file and its bytes for good,' naming the exact verb and resource. It also distinguishes deletion from sibling file operations by detailing version and retention constraints, so an agent understands this is the permanent-delete tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states when deletion is blocked (retention_locked, not_head_version) and how to complete a staged deletion with confirmed=true and the same idempotencyKey. It does not explicitly name alternative tools like files_set_retention or archive_item, but it gives enough context for an agent to decide whether and how to call delete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_get_contentBRead-only
Read a stored file back as base64, having re-hashed the bytes and compared them against the stored sha256 first. A mismatch returns integrity_mismatch with both hashes and NO bytes, because serving content the checksum disowns is worse than refusing; a missing blob returns content_not_found. This is the OR 958f Abs. 3 readability half. Needs the file-content right on top of the master-data read: a read-only role may list the filing and not download it.
| Name | Required | Description | Default |
|---|---|---|---|
| fileId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses important behavioral traits: it re-hashes bytes and compares against stored sha256, returns integrity_mismatch with both hashes on mismatch, and content_not_found if the blob is missing. It also explains the rationale for refusing to serve mismatched content and mentions permission nuances. This is substantial additional context, though the cryptic 'OR 958f' phrase adds confusion.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded ('Read a stored file back as base64'), but the description becomes verbose with conditional error behaviors and cryptic references. Phrases like 'This is the OR 958f Abs. 3 readability half' and 'Needs the file-content right on top of the master-data read' are unclear and add noise without clear value. The structure could be tightened by focusing on the action, parameters, and expected outputs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description should explain what the success response contains beyond base64, but it only mentions error cases. It also fails to describe how workspaceId and fileId are used or validated. The cryptic legal/regulation reference does not aid understanding. For a tool with two parameters and no schema descriptions, this is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has two required string parameters (workspaceId and fileId) with no descriptions (schema coverage 0%). The description does not mention either parameter, leaving the agent without clues about what fileId refers to (e.g., a file identifier) or how workspaceId scopes the operation. This is a critical gap for a tool with no schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads a stored file and returns it as base64, with a specific verb ('Read') and resource ('stored file'). It also mentions the integrity check and error conditions, making it distinct from sibling file tools like files_upload, files_search, or files_delete, which have different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus alternatives. It includes a cryptic reference to 'OR 958f Abs. 3 readability half' and a note about read-only roles, but no clear guidance on when to choose this over other file operations. The agent must infer that this is the primary content-retrieval tool, but the description lacks explicit routing or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_linkA
Attach a stored file to any record TILL knows (the OP3 linkEntity): entityKind is validated against the shared entity registry and entityId must exist in this workspace. Linking to POSTED accounting evidence (a document whose issue posted an entry, a payment, a posted journal entry) derives the OR 958f ten-year lock from the END of the fiscal year that RECORD belongs to, so a future-dated entry is kept longer and a backdated one is kept from its own year. A link to a record that has not posted derives nothing until it posts (D63): a draft, and equally an issued quote or order, which carry a number but no booking. The floor attaches at the posting moment and is then permanent, and a never-posted draft that is deleted leaves no lock behind. The derived date is remembered separately, so it survives a later manual extension and a re-link. Re-linking replaces the link. Requires whatever writing the target itself requires.
| Name | Required | Description | Default |
|---|---|---|---|
| fileId | Yes | ||
| entityId | Yes | ||
| entityKind | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does an excellent job. It discloses validation requirements, the derived retention lock derivation based on fiscal year, the behavior for non-posted records, the floor that becomes permanent at posting, re-linking replacement, survival of derived dates, and the requirement for write permissions. These are rich behavioral details beyond what a simple annotation could provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is lengthy but dense with necessary detail. The opening sentence clearly states the core action, and subsequent sentences each add distinct behavioral information. No filler or repetition. It is front-loaded and organized logically, though it could be trimmed slightly without losing clarity, but the complexity justifies the length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity and lack of annotations and output schema, the description should cover all critical aspects. It covers validation, retention behavior, and write requirements, but omits what the tool returns (if anything), error conditions beyond validation, and the purpose of idempotencyKey. Some gaps remain for an agent to invoke it correctly without further context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero descriptions (coverage 0%), so the description must compensate. It adds meaning for entityKind (validated against shared entity registry) and entityId (must exist in workspace), and explains consequences for linking to certain record types. However, it does not define workspaceId, fileId, or idempotencyKey semantics, nor does it enumerate valid entityKind values. Partial compensation, but not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the action ('Attach a stored file to any record TILL knows') and the resource (a file to a record). It also distinguishes itself from sibling file tools by focusing on the linking behavior and its domain-specific consequences, not just a generic attach. The purpose is unambiguous and specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for attaching files to records but does not explicitly state when to use it over alternatives like files_list_linked or files_upload. It gives context about validation and derived lock behavior, which hints at applicability, but there is no explicit when-not or alternative routing. The guidance is implied rather than direct.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_list_linkedARead-only
Every file attached to one record (the OP3 listLinked), newest first. Current versions only; includeVersions nests each file's history underneath it rather than listing a superseded copy as a peer.
| Name | Required | Description | Default |
|---|---|---|---|
| entityId | Yes | ||
| entityKind | Yes | ||
| workspaceId | Yes | ||
| includeVersions | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the operation as read-only, and the description adds useful behavior beyond that: newest-first ordering, current-versions-only semantics, and the fact that includeVersions nests history rather than listing superseded copies as peers. No contradiction with the readOnlyHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, each earning its place. The primary behavior is front-loaded, and the second sentence provides a precise, compact clarification of the subtle version behavior without excess wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema, the description covers the essential selection and version-nesting behavior. It lacks details about pagination or the meaning of entityKind, but the core call context is mostly sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description should carry more semantic weight. It explains includeVersions clearly, and 'one record' loosely maps to entityId, but workspaceId and entityKind are left unexplained, including what values entityKind accepts or how they scope the result.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: listing files attached to one record, and clarifies version and ordering behavior. It is clear on its own but does not explicitly distinguish itself from sibling file tools like files_search or files_get_content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The scope 'attached to one record' implies when the tool is appropriate, and the version behavior implies when includeVersions should be used. However, it does not explicitly state when to prefer an alternative such as files_search, files_link, or files_get_content.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_new_versionA
Add a version that supersedes the current one, keeping both. The chain is linear and append-only (OR 958f): the prior version stays readable and stays retained, and superseding a version something already supersedes is refused with not_head_version. Folder, title, tags, the entity link and the retention date all carry forward.
| Name | Required | Description | Default |
|---|---|---|---|
| mime | No | ||
| fileId | Yes | ||
| filename | No | ||
| workspaceId | Yes | ||
| contentBase64 | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly states that the chain is linear and append-only, the prior version remains readable and retained, superseding a non-head version is refused, and folder/title/tags/entity link/retention date carry forward. This exceeds what the schema reveals and gives the agent a clear model of side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is relatively compact and well-structured, leading with the core action, then adding constraints and metadata inheritance. The internal reference '(OR 958f)' is cryptic and of marginal value, but does not detract significantly from overall efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers behavior, constraints, and metadata carry-forward, which is substantial for a version-add operation. However, there is no output schema, and the description does not mention the response shape or return value, nor does it elaborate on required parameters. These are notable omissions, though the description is still quite thorough.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does mention that folder, title, tags, entity link, and retention date carry forward, which implies those fields need not be provided. However, it does not explain the required parameters (workspaceId, fileId, contentBase64) or optional ones like idempotencyKey, mime, and filename, leaving a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: adding a version that supersedes the current one while keeping both. This specific verb-resource pairing distinguishes it from siblings like files_update or files_upload, and the append-only, keep-both behavior further disambiguates it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context is provided for when to use this tool (when a new version is needed that supersedes the current one while preserving old versions). It implies the head-version constraint via the not_head_version error, but does not explicitly name alternative tools or articulate when not to use them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_searchARead-only
Find stored files by text over title, filename and tags, and by folder, tag, mime or linked entity kind. Every whitespace-separated word in q must match, so a second word narrows. Current versions only, newest first, with a truncation flag past the documented ceiling.
| Name | Required | Description | Default |
|---|---|---|---|
| q | No | ||
| tag | No | ||
| mime | No | ||
| folderId | No | ||
| entityKind | No | ||
| workspaceId | Yes | ||
| includeVersions | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses non-obvious behaviors: q is a multi-term AND match, results are current versions only, ordering is newest first, and a truncation flag appears past the documented ceiling. This gives the agent useful runtime expectations beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences cover search scope, query semantics, ordering, and truncation behavior. There is no filler or repetition of the annotations, and the most important usage detail is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with no output schema and 0% schema description coverage, the description covers the main facets the agent needs: searchable fields, q matching behavior, version scope, ordering, and truncation signaling. It falls slightly short only on exact return shape and how filters combine with q.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden for parameters. It explains q semantics and maps folder, tag, mime, and entityKind to search criteria, which is valuable. However, workspaceId and especially includeVersions are not explicitly clarified against 'Current versions only', leaving ambiguity for those parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific operation (searching stored files) over a defined resource type and enumerates the searchable dimensions: title, filename, tags, folder, tag, mime, and linked entity kind. It also adds behavioral scope ('Current versions only'), so the purpose is unambiguous, though it does not explicitly differentiate from sibling search/list tools like files_list or search_global.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives practical query guidance ('Every whitespace-separated word in q must match, so a second word narrows') and scoping ('Current versions only'), which implies when this tool is appropriate. However, it never names alternative tools or states when to prefer files_list, files_list_linked, or search_global, so exclusions must be inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_set_retentionA
Set or extend how long a file must be kept (Aufbewahrung bis, ISO YYYY-MM-DD). Extending is always allowed; a date below the OR 958f floor is refused with retention_below_statutory and the floor it computed. The floor is remembered separately from the date, so extending a statutory retention by hand cannot erase it and a later re-link to a non-accounting record cannot release the file. A date set by hand is recorded as manual provenance, never as gesetzlich.
| Name | Required | Description | Default |
|---|---|---|---|
| fileId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No | ||
| retentionUntil | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure and does so thoroughly. It reveals the always-allowed extension rule, the refusal with a specific error code (retention_below_statutory), the separately remembered statutory floor, and provenance tracking, and the distinct manual vs. gesetzlich classification. This goes far beyond a simple action statement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three dense sentences with no filler. The primary action and format are front-loaded, and each subsequent sentence provides critical behavioral edge-case information that earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderately high complexity, the complete absence of annotations, and no output schema, the description covers the main action, input format, error case, and subtle statutory-floor semantics. An agent has what it needs to invoke the tool correctly and understand the outcome boundaries, including the refusal condition and provenance handling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Since schema description coverage is 0%, the description must compensate. It adds essential meaning for retentionUntil (ISO YYYY-MM-DD, floor refusal behavior, manual provenance) but does not describe the conventional workspaceId, fileId, or idempotencyKey parameters. Still, it significantly enriches the most complex parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Set or extend how long a file must be kept,' clearly distinguishing this from sibling file tools like files_delete, files_link, or files_update. The retention-specific scope is unambiguous, and the format and error semantics further pin down the exact purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context (set or extend retention, extending always allowed) but does not explicitly state when to use this tool versus alternatives such as files_delete or files_update. Since no sibling tool appears to handle retention, the intended usage is implied rather than explicitly contrasted.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_updateA
Patch a stored file's filing metadata: title, tags, folder, or the pendingDelete flag an agent-staged deletion set. Never its bytes: content is corrected by a new version, never by an edit, so sha256, size and version have no path through this verb.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | Yes | ||
| fileId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of disclosure. It transparently states that bytes are never modified, that sha256, size, and version cannot be changed through this verb, and that pendingDelete can be set. This gives clear immutability caveats and side-effect expectations. It does not mention response format, error cases, or whether patching triggers version bumps, but the core behavioral scope is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero fluff. The first sentence states the primary purpose and allowed fields; the second adds the critical exclusion. Both are front-loaded and each earns its place, making the description highly scannable for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the nested patch object, few parameters, and no output schema, the description effectively covers the operation's scope. An agent can infer how to construct the call from the field list and the exclusion of content fields. Missing details like idempotencyKey semantics or response shape are minor because they are standard patterns, and the description's clarity about the patch payload is sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It does by explaining what the patch object may contain (title, tags, folder, pendingDelete) and explicitly lists what it cannot (sha256, size, version). This gives meaning to the most important parameter. The other parameters (workspaceId, fileId, idempotencyKey) are not explained, but they are common and inferable from naming; the description still adds enough value for the critical param.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (patch) and resource (stored file's filing metadata), and enumerates the exact mutable fields: title, tags, folder, and pendingDelete. It explicitly disclaims any content modification, sharply distinguishing it from files_new_version and other content-altering tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear boundary: use this to patch metadata, never for bytes. It points to 'a new version' as the correct path for content changes, effectively naming the alternative without explicitly listing it. It does not list other exclusions (e.g., deletion via files_delete), but the scope is unambiguous enough for an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_uploadA
Store a file on this device: the bytes arrive base64-encoded in contentBase64, are hashed with sha256, and are kept content-addressed inside the ledger so nothing leaves the machine. Optionally filed into a folder and tagged. Refuses a payload that is not base64 (file_unreadable) or larger than 25 MiB (file_too_large), and never trusts a caller-supplied hash.
| Name | Required | Description | Default |
|---|---|---|---|
| mime | No | ||
| tags | No | ||
| title | No | ||
| filename | No | ||
| folderId | No | ||
| workspaceId | Yes | ||
| contentBase64 | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations available, the description carries the full burden and delivers: sha256 hashing, content-addressed ledger storage, no machine egress, refusal modes with error codes, and no trust in caller-supplied hashes. It clearly discloses side effects and constraints without contradicting any annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the core action, mechanism, and constraints. Every sentence adds distinct value with no filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-annotation, no-output-schema upload tool, the description covers core behavior, constraints, storage semantics, and error conditions, plus optional filing and tagging. It does not state what is returned after a successful upload or how the stored file is subsequently referenced, and workspaceId semantics are left entirely to the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain contentBase64's base64 encoding and the 25 MiB limit, and it names optional folder/tag filing. However, required workspaceId and optional title, filename, mime, and idempotencyKey are not explained, leaving several parameters under-documented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a concrete operation (store a file), a target (this device/ledger), and a specific encoding/hashing model. It is easy to tell apart from chunked or versioned file siblings because it describes a single base64 payload with a size cap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly conveys the main use case: supply a full base64 payload, optionally filing into a folder and tagging, with a 25 MiB limit. However, it never names alternatives like files_upload_begin/chunk/commit for larger files or files_new_version for replacing content, so routing guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_upload_beginA
Open a chunked upload for a file too large for the 25 MiB single-call bound (G18 US-G18.4): declare its name, mediaType, sizeBytes and intent (migration_source), and receive an uploadId plus the per-chunk byte ceiling. The bytes then arrive through files_upload_chunk and the blob is minted only by files_upload_commit once the accumulated sha256 is verified. The migration-class ceiling is 500 MB; an incomplete session expires after 24 h and never leaves a stored file behind.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| intent | No | ||
| mediaType | No | ||
| sizeBytes | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure and does so thoroughly. It reveals that no stored file is left behind until commit, sessions expire after 24 hours, and the blob is only minted after sha256 verification, which are critical non-obvious behaviors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences carry no filler; the trigger condition, workflow steps, limits, and expiry are all front-loaded in logical order. Even the specification reference (G18 US-G18.4) earns its place as a precise policy pointer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a stateful upload initiation with no output schema, the description supplies the return value (uploadId and per-chunk ceiling), the companion calls, verification requirement, size ceiling, and expiry behavior. The only material gap is that required workspaceId and the optional idempotencyKey are not explained, which an agent would need to construct a fully correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides zero descriptions, so the text meaningfully adds semantics for name, mediaType, sizeBytes, and intent, including the intent value migration_source. However, it omits workspaceId, despite being a required parameter, and does not explain idempotencyKey, so not all parameters are covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource, 'Open a chunked upload', and immediately distinguishes this start step from companion tools by naming files_upload_chunk and files_upload_commit. It also states the concrete trigger condition (files too large for the 25 MiB single-call bound), making its role unambiguous among hundreds of sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use this tool: for files exceeding the 25 MiB single-call limit, up to the 500 MB migration-class ceiling. It also names the exact follow-on tools (files_upload_chunk, files_upload_commit) and explains the state transition, so an agent can orchestrate the full protocol.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_upload_chunkA
Append one ordered chunk (base64, each within the 25 MiB bound) to an open upload session. Chunks arrive in seq order from 0; a repeated seq is an idempotent no-op and a gap is refused with chunk_out_of_order. The running total may not exceed the declared sizeBytes.
| Name | Required | Description | Default |
|---|---|---|---|
| seq | Yes | ||
| uploadId | Yes | ||
| workspaceId | Yes | ||
| contentBase64 | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the burden. It discloses statefulness (open upload session), ordering semantics, the exact failure mode for gaps (chunk_out_of_order), idempotency for repeated seq, and the overall size cap relative to declared sizeBytes. It doesn't state auth requirements or what happens on final chunk, but that's covered by commit sibling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense, front-loaded sentences with no filler. Every clause adds constraint or behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a chunked transfer tool, the protocol-critical info (order, size cap, gap refusal) is all present, and the commit/begin siblings supply the surrounding workflow. Minor gap: does not state what a successful append returns (schema has no output schema, so not required).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It gives semantics for seq (ordered), contentBase64 (base64, 25 MiB bound), and sizeBytes cap at session level, but does not explain workspaceId, uploadId, or how idempotencyKey interacts beyond the implied no-op.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (append), resource (open upload session), constraint (chunks in seq order within 25 MiB), and semantics (idempotent no-op, gap refused). It fully distinguishes from sibling files_upload_begin/files_upload_commit/files_upload/files_new_version.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes ordering requirement and the idempotent/out-of-order behaviors that govern when a call is valid. It doesn't name alternatives (files_upload versus files_upload_chunk), but the protocol context is clear enough to select this tool during a chunked upload sequence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
files_upload_commitA
Close an upload session and mint the blob, verifying the accumulated sha256 against the caller-declared one: a mismatch refuses with source_integrity_mismatch and stores nothing. A blob over the single-call bound is kept in segments and is read only through the streaming reader (files_get_content refuses it with file_too_large_use_stream). Idempotent: a replayed commit returns the file it already minted.
| Name | Required | Description | Default |
|---|---|---|---|
| sha256 | Yes | ||
| uploadId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly states what happens on hash mismatch (refuses with source_integrity_mismatch and stores nothing), how large blobs are handled (segmented, only readable via streaming), and that it is idempotent (replayed commit returns the already-minted file). This is comprehensive and directly informs the agent of side effects and edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, dense with useful information, and front-loads the primary action. Every sentence earns its place: the first defines the main operation and error case, the second covers large-blob handling and the related alternative tool, and the third adds idempotency. There is no fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description covers the key behavioral aspects: error handling, large-blob implications, and idempotency. It does not explicitly describe the return value for a successful commit (only the idempotent case), nor does it state prerequisites (e.g., that an upload must have been started). However, the complexity is moderate, and the description provides sufficient context for an agent to invoke it correctly in most scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. It explains sha256 (the caller-declared hash to verify against) and uploadId (the session to close), but does not clarify workspaceId or idempotencyKey. While workspaceId is likely a standard workspace identifier and idempotencyKey is optional, the description leaves these unexplained. Given the lack of schema descriptions, more explicit parameter guidance would be beneficial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: closing an upload session and minting a blob, with a specific verification step (sha256 check). It distinguishes this from sibling tools like files_upload_begin (which starts the upload) and files_upload_chunk (which uploads data) by naming the action as a commit/close operation. The verb 'Close' and resource 'upload session' are precise, and the integrity check is a distinctive feature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is the final step in an upload flow (closing a session), and it explicitly mentions an alternative for reading large blobs via the streaming reader (files_get_content). However, it does not explicitly state the order of operations (e.g., 'call after files_upload_begin and files_upload_chunk') nor explicitly say when not to use it. The context is strong but could be more direct about the upload sequence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
flag_entryA
Flag a posted journal entry as questioned (beanstandet) with a reason: review metadata only, the posted rows stay untouched. The entry shows as flagged in review_status until someone approves it; the books are corrected by reversal, never by edit.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| entryId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden, and it does well: it discloses that only review metadata is affected, posted rows stay untouched, the review_status reflects the flag until approval, and corrections come from reversal. It does not mention idempotency behavior, permissions, or exact response details, but the core side effects and states are clearly conveyed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler at all. The primary action and the boundary condition are front-loaded, followed by lifecycle context that an agent needs for accurate invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains the operation's scope, lifecycle, and key side effect, which is substantial given the absence of annotations. It leaves gaps in parameter semantics, particularly idempotencyKey and return/error behavior, and does not mention prerequisites like workspace membership. These omissions matter for a write operation with an idempotencyKey.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only explains the 'reason' parameter. workspaceId and entryId are left to inference, and idempotencyKey—critical for a potentially non-idempotent operation—is completely unexplained. The added meaning over the bare schema is minimal.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Flag') and resource ('posted journal entry') with the exact intent ('questioned (beanstandet)') and a reason. It distinguishes itself from sibling tools by explicitly noting it only alters review metadata and leaves posted rows untouched, which separates it from post_entry, reverse_entry, approve_entry, and edit-like actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool: to mark an entry as questioned without changing the underlying posted rows. It also clarifies alternatives by noting that corrections are done by reversal, not edit, and that the flag persists until approval, implying approve_entry is a separate step. However, it does not explicitly name sibling tools or list exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
folders_deleteA
Delete a filing folder, and only when it holds nothing at all. A folder containing files or child folders is refused with folder_not_empty and both counts: deletion never cascades, because a cascade could erase a retained record nested inside it without ever consulting its retention lock.
| Name | Required | Description | Default |
|---|---|---|---|
| folderId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully owns behavioral disclosure. It explicitly reveals that the operation is refused for non-empty folders with folder_not_empty and includes both counts, and that deletion never cascades. It even explains the rationale (avoiding erasure of retained records without checking retention locks), giving agents deep insight into the safety behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with the core behavior front-loaded. The second sentence adds essential safety detail (refusal behavior and non-cascade) rather than fluff, and the rationale about retention locks justifies the length. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple delete operation with a small schema and no output schema, the description covers the key contextual facts: success condition, failure condition, error code, and non-cascade behavior. It doesn't mention what a successful call returns or how idempotencyKey is used, but these are minor given the tool's simplicity and the absence of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the three parameters (workspaceId, folderId, idempotencyKey), and the description adds no parameter-specific meaning. While folderId is self-evident from the tool name and workspaceId is a common context parameter, the description does not compensate for the absence of schema documentation, leaving idempotencyKey's purpose entirely undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('delete a filing folder') and a critical condition ('only when it holds nothing at all'). It distinguishes this from folder creation/upsert and file deletion by focusing on the empty-folder deletion scenario. The failure mode (folder_not_empty) is explicitly named, further clarifying the tool's purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when this tool should be used ('only when it holds nothing at all') and the consequence if that condition isn't met. It does not explicitly reference alternative tools (e.g., folders_upsert for non-empty folders), but the 'only when empty' phrasing provides clear usage context and effectively excludes non-empty cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
folders_listARead-only
The filing tree as a flat list in path order, which is already tree order: each folder carries its depth, its direct file count, its child count, and whether it is deletable, so a client renders the rail from one read and never re-derives the delete rule differently from the engine.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true covering the safety profile, the description adds substantial behavioral context: the result is a flat list ordered by path (already tree order), and each folder carries depth, direct file count, child count, and a deletable flag. It further discloses an invariant — the client never re-derives the delete rule differently from the engine — which is valuable for callers relying on consistent behavior. No contradiction with the annotation exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense but efficient sentence; every clause contributes (ordering, per-folder fields, client use case, delete-rule invariant). The rationale clause is slightly redundant with the field list but earns its place by explaining intent, and the core semantics are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read with no output schema, the description compensates well by spelling out the return shape, ordering, and computed fields. It omits edge-case behavior (empty tree, invalid workspaceId, potential large-result handling), but for a low-complexity tool these are acceptable gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description was expected to compensate, but it never mentions workspaceId. However, the single parameter is well-named, required, and typed as a plain string, so its meaning is effectively self-evident; the gap is minor rather than harmful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific resource (the filing tree) and a specific output form ('flat list in path order'), so an agent understands this is the read/list operation for folders. It does not use an explicit imperative verb, but the semantics are unmistakable and the content clearly differentiates it from the mutation siblings folders_upsert and folders_delete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The closing clause ('so a client renders the rail from one read') implies when this tool is appropriate — when a client needs the whole folder tree for rendering — but this is framed as a design rationale rather than explicit usage guidance. No alternatives are named and no when-not-to-use conditions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
folders_upsertA
Create a filing folder, or rename and re-parent an existing one by passing its folderId. The full path is materialised and re-materialised for every descendant in the same transaction, sibling names are unique, and a folder can never become a descendant of itself (folder_cycle).
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| folderId | No | ||
| parentId | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing behavioral traits. It goes beyond the basic action by explaining that the full path is materialised for descendants transactionally, sibling names are unique, and a folder cannot become a descendant of itself. These are significant invariants and side effects that an agent needs to know, and they are clearly disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, with the primary action front-loaded. It efficiently states the purpose and then adds critical constraints without excess. Every sentence contributes meaningful information, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters and no output schema or annotations, the description is incomplete. It fails to clarify the required parameters for create versus update operations, does not explain workspaceId or idempotencyKey, and provides no information about the response format or error behavior. While it covers some behavioral constraints, the missing parameter usage details and output expectations make it insufficient for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage, so the description must compensate for parameter meaning. It only mentions folderId as the identifier for existing folders and implies name and parentId for creation/update, but it does not explain workspaceId (required) or idempotencyKey. The semantics for create vs update parameters are vague, leaving the agent uncertain about which parameters are necessary in each scenario.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: create a filing folder or rename/re-parent an existing one by passing its folderId. It names the specific verb and resource, and distinguishes from siblings like folders_list and folders_delete by indicating it handles create/update operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context that this tool is for creating or updating folders, which implies when to use it. However, it does not explicitly mention alternatives or when not to use it, such as stating to use folders_list for listing or folders_delete for deleting. The guidance is implied rather than explicit, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
forecast_revenueARead-only
Die Umsatzprognose (revenue forecast) per month over a 1-24 month horizon, three labelled components with each expected Rappen in exactly ONE: weightedOpenMinor (open deals, probability-weighted, overdue or dateless closes in the first period), wonUninvoicedMinor (won deals with no A11 invoice reachable over the document source chain, at 100 percent), and openQuotesMinor (sent, unexpired C02 quotes with NO deal on either link side, at 100 percent of their base-currency total; a deal-linked quote is represented by its deal, so nothing double-counts). A foreign-currency quote has no stored CHF base and lands in excluded[] with needs_fx_rate instead of a silently wrong total. A projection (Prognose, keine Buchhaltungszahl): nothing here is posted revenue.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| horizonMonths | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description adds substantial behavioral detail: three mutually exclusive components, probability weighting, no double-counting, foreign-currency exclusion with needs_fx_rate, and the clear statement that output is not posted revenue. This is far more than annotations alone provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded and every clause adds non-redundant detail. However, the text is a dense run-on sentence with awkward phrasing like 'each expected Rappen in exactly ONE' and several unexplained abbreviations, so structure could be cleaner.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of an output schema, the description covers the key behavioral context well: components, exclusions, FX handling, and double-counting prevention. It falls slightly short on explicitly defining the full return shape and the workspaceId parameter, but for a two-parameter forecast tool it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden. It clarifies horizonMonths as a 1-24 month horizon, but it never mentions workspaceId, and most of the description explains output behavior rather than parameter meaning. Partial compensation only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource (revenue forecast) and the computation (per month over a 1-24 month horizon), with specific output components. It does not explicitly differentiate from sibling forecasting tools such as forecast_sales_kpis or forecast_weighted_pipeline, but the level of detail makes the purpose clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes an explicit when-not statement ('nothing here is posted revenue') and clarifies this is a projection rather than accounting data. However, it does not mention alternative forecasting tools or state selection criteria for when to use this tool instead of siblings, so usage guidance is mostly implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
forecast_sales_kpisARead-only
Verkaufskennzahlen (sales KPIs) over deals CLOSED in the from/to window: conversionRateBp = won over (won plus lost) in integer basis points, avgDealSizeMinor = round-once mean of won base values, avgCycleDays = whole-day mean from deal creation to the OP5 won-timestamp. Zero closed deals answers null KPIs with sample:0, never a fake 0 percent an agent could misread. Management figures over pipeline estimates, not accounting figures.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | ||
| from | Yes | ||
| pipelineId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the readOnlyHint annotation by disclosing the zero-deals edge case (returns null KPIs with sample:0, never a fake 0 percent) and specifying calculation details (round-once mean, whole-day mean, OP5 won-timestamp). This is exactly the kind of behavioral context an agent needs, and it does not contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence that packs purpose, formulas, edge case, and a caveat. It is front-loaded with the main purpose and contains no fluff, though the length and semicolon-heavy structure could be slightly more scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only KPI tool, it covers the essential output metrics, the zero-deals edge case, and the distinction from accounting figures. However, without an output schema, it does not describe the return envelope (e.g., JSON structure, units for avgCycleDays) or parameter formats like date strings, leaving some context incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameters. It only clarifies the 'from/to window' and does not mention workspaceId or pipelineId, which remain undocumented. The KPI formulas are explained but not mapped to parameters, leaving a significant gap for a 4-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool computes sales KPIs over deals CLOSED in a time window, and explicitly names each metric (conversionRateBp, avgDealSizeMinor, avgCycleDays) with formulas. It also distinguishes itself from pipeline estimates via 'Management figures over pipeline estimates, not accounting figures.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: use for closed-deal management KPIs, not accounting figures. It implicitly contrasts with pipeline estimates ('over pipeline estimates'), which suggests when this tool is appropriate vs. other forecast tools. However, it does not explicitly name sibling tools like forecast_weighted_pipeline or forecast_vs_actual.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
forecast_vs_actualARead-only
Prognose vs. Ist (forecast vs actual) for one period (YYYY, YYYY-MM or YYYY-Qn): actualRevenueMinor is the A08 Erfolgsrechnung's Nettoerlöse subtotal (POSTED revenue, net of VAT, the only accounting figure in this capability) against wonInPeriodMinor (won-deal base value with the OP5 close date in the period), with the delta and the two named gap buckets: wonNotInvoicedMinor (revenue still to come) and invoicedWithoutDealMinor (posted invoices whose source chain reaches no deal, at their net subtotal). A period with neither answers zeros with sample:0. Stores no snapshot (P5): scheduled history belongs to F01.
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint, the description discloses meaningful behavior: what actualRevenueMinor and wonInPeriodMinor mean at source level, the zero/zero-blank period response behavior with sample:0, and the no-snapshot policy with F01 ownership. This is exactly the kind of context an agent needs and it contradicts no annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
There is no filler; every clause adds useful information, and the core comparison is front-loaded. The density and run-on style could be improved with structuring, but it remains efficient for the complexity it covers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description compensates by naming the returned concepts: delta, wonNotInvoicedMinor, invoicedWithoutDealMinor, and the sample:0 edge case. It does not describe the response envelope or error behavior, but given the read-only scope and detailed semantics, it is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate, and it does for the period parameter by defining the accepted formats YYYY, YYYY-MM and YYYY-Qn and clarifying what period means. workspaceId remains implicit and self-evident from the parameter name, which prevents a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states exactly what the tool computes: a forecast-vs-actual comparison for one period, comparing actualRevenueMinor against wonInPeriodMinor, the delta, and two named gap buckets. It names precise source definitions and the allowed period format, which makes it clearly distinguishable from the forecasting and reporting siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context on when this tool applies: a single period with specific formats, and it explicitly notes that no snapshot is stored and scheduled history belongs to F01, which routes agents away for historical needs. It does not explicitly name alternative forecast tools like forecast_revenue or forecast_weighted_pipeline, so it is not a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
forecast_weighted_pipelineBRead-only
Die gewichtete Pipeline (weighted pipeline): every OPEN deal contributes round-once(valueBaseMinor times probability over 100), C01's own formula, grouped by 'stage' | 'month' | 'quarter' of expectedCloseOn or by any confirmed select/multiselect custom-field key on deal (rows re-bucket, totals never move: totalWeightedMinor is identical for every valid groupBy). Sums are integer Rappen over the stored CHF base (never re-converted); deals without a close date bucket under 'none', never dropped. An empty pipeline answers rows:[] and zero totals, not an error. This is a PROJECTION over estimates (Prognose, keine Buchhaltungszahl), not an accounting figure: revenue exists when A11 issues an invoice.
| Name | Required | Description | Default |
|---|---|---|---|
| groupBy | No | ||
| pipelineId | No | ||
| workspaceId | Yes | ||
| horizonMonths | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses several behaviors beyond the readOnlyHint annotation: the rounding formula, integer Rappen over CHF base, handling of deals without close dates, empty pipeline response, and the projection-vs-accounting distinction. It does not mention errors or rate limits but covers key edge cases well.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph with many clauses. It front-loads the core formula but is somewhat unstructured, mixing multiple concepts (rounding, grouping, edge cases) without clear separation. It is informative but not optimally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the formula, grouping, edge cases, and output shape for empty results, but it omits details about the required workspaceId, pipelineId selection, and horizonMonths semantics. Without an output schema, it also does not describe the response structure beyond rows and totals. It is complete for the main use case but leaves gaps for parameters and response details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must explain all parameters. It explains groupBy (stage/month/quarter or custom field) but does not describe workspaceId, pipelineId, or horizonMonths. This leaves significant ambiguity for the required and other parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool computes a weighted pipeline forecast with a specific formula, grouping options, and edge cases. It is distinct from forecast_revenue and forecast_sales_kpis, though it does not explicitly name alternatives, so it is clear but not fully differentiated from sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus other forecasting tools such as forecast_revenue or forecast_vs_actual. The description explains behavior but does not provide context on selection criteria or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fx_revaluationARead-only
Revalue every open foreign-currency monetary position (A19 FC bank balances, A16/A17 open FC debtor/creditor items) at the closing rate for periodEnd (OR Art. 960a: Bilanzstichtagskurs). Returns each position with its foreign-currency amount, book CHF, closing rate, revalued CHF and the UNREALISED diff, the per-currency subtotals, the net total, and any currency that still needs a rate. A read model: nothing is posted. Negative total is a net unrealised loss.
| Name | Required | Description | Default |
|---|---|---|---|
| periodEnd | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces this with 'nothing is posted'. It adds useful behavioral detail beyond annotations: the full return structure, per-currency subtotals, the unrealised-diff interpretation, and the signal that some currencies may still need a rate. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause adds value: scope, valuation basis, return fields, safety, and sign interpretation. The return enumeration is long and slightly run-on, but it is justified because there is no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex financial tool with no output schema, the description covers the purpose, parameter semantics for periodEnd, the full read-model return, and the meaning of a negative total. The main missing context is an explicit pointer to post_fx_revaluation and whether rates must be recorded beforehand, but those are secondary for invoking a read-only preview.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry parameter meaning. It meaningfully explains `periodEnd` as the closing-rate date and legal basis, but it does not explain `workspaceId` at all. This is a partial compensation rather than full coverage of the required parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Revalue'), names the exact scope (A19 FC bank balances, A16/A17 open debtor/creditor items), and specifies the valuation basis ('closing rate for periodEnd', Art. 960a). It also explicitly frames itself as a read model, which distinguishes it from siblings like post_fx_revaluation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use it: 'A read model: nothing is posted' signals this is for previewing rather than posting a revaluation. However, it does not name the alternative (post_fx_revaluation) or state explicit conditions such as 'use this before posting'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fx_revaluation_reverseA
Revert a posted FX revaluation run (D129 Q2): books the mirror of the revaluation entry dated the period end (source fx) plus its own next-day reversal, atomically, so every position and account 6949 net to zero on both dates and nothing is edited. runId is the run post_fx_revaluation returned. Refuses not_posted for a run that posted nothing, already_reversed, later_run_exists naming the later period end whose run still stands (revert newest first), and period_locked for a locked period. Idempotent on idempotencyKey. CONSEQUENCE: Books the mirror of the period-end currency revaluation and its own next-day reversal; the original entries stay on record.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries the behavioral burden. It discloses atomicity, that the mirror entry and its own next-day reversal are booked, that positions and account 6949 net to zero on both dates, and that original entries stay on record. There is no contradiction with annotations since none exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded: the first sentence states purpose and mechanism, then parameter origin, refusal conditions, and idempotency. The CONSEQUENCE line at the end repeats what was already stated, adding slight redundancy, but every other clause earns its place for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description thoroughly covers behavior, error conditions, and parameter origins. It lacks an explicit return value description and omits workspaceId explanation, but the core invocation requirements are fully specified and an agent can call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains runId as the run returned by post_fx_revaluation and identifies idempotencyKey as the idempotency key, but it does not explain workspaceId at all. This is only partial compensation for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reverts a posted FX revaluation run and details the exact mechanism (books the mirror plus its own next-day reversal). It references the sibling post_fx_revaluation by saying runId is the run that tool returned, which differentiates it from generic reversal tools like reverse_entry.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit preconditions and refusal conditions: runId must come from post_fx_revaluation, and it refuses for not_posted, already_reversed, later_run_exists (implying 'revert newest first'), and period_locked. It also states idempotency on idempotencyKey. It does not explicitly name alternative tools for other reversal scenarios, so it lacks full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
general_ledgerARead-only
Kontoblatt (general ledger) for ONE account: the opening carry, every posted line in the period in date order with its running balance, and the closing balance. Each line carries the entryId to drill into with get_entry. Debit-positive like the Saldenbilanz, with naturalSide naming which way the account is expected to lean. Drafts never appear. groupBy accepts 'kmu' today and refuses anything else.
| Name | Required | Description | Default |
|---|---|---|---|
| groupBy | No | ||
| accountId | Yes | ||
| periodEnd | Yes | ||
| periodStart | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true, so the safe read nature is already known. The description adds valuable behavioral details: drafts never appear, debit-positive convention, naturalSide naming expectation, and groupBy restriction to 'kmu'. These go beyond annotations and help the agent understand the response structure and constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured, front-loading the core purpose (Kontoblatt for ONE account) and what it returns, then adding key behavioral notes. Each sentence adds value; there is no filler or repetition. The groupBy restriction is clearly stated at the end without unnecessary elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no output schema, the description adequately explains the return structure: opening carry, lines with running balance, closing balance. It also explains the drill-down capability via entryId, which covers a key use case. The groupBy limitation is disclosed, and the necessary parameters are clear. This is complete for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies the mandatory parameters (workspaceId, accountId, periodStart, periodEnd) by contextualizing them (e.g., the period range for the ledger). It also explains the groupBy parameter by stating it only accepts 'kmu', which is not evident from the schema. This adds substantial meaning beyond the bare parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is a 'Kontoblatt (general ledger) for ONE account' and specifies the exact content: opening carry, posted lines with running balance, closing balance. It distinguishes itself from a general trial balance or journal by emphasizing it is per-account and includes running balances, which differentiates it from siblings like list_journal or trial_balance.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly indicates when to use this tool: when you need the detailed ledger for a single account, as opposed to a trial balance or journal. It does not explicitly name alternatives or provide exclusion criteria, but the emphasis on 'ONE account' and the mention of the entryId for drill-down with get_entry gives clear context for appropriate use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
generate_pain001A
Build (or rebuild) the batch's pain.001.001.09 XML, validate it, and flip the batch from draft to generated on first success; regenerating an already-generated batch reproduces byte-identical output (nothing about it changes after creation). Only a valid file is ever returned. NEVER transmits: always answers transmitted:false, with reason 'use_payment_batch_transmit' when an A33 EBICS channel routes the batch's debtor account (transmit is A33's payment_batch_transmit, P8-gated) or 'no_channel' otherwise (the file-download path is then the floor).
| Name | Required | Description | Default |
|---|---|---|---|
| batchId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it is unusually explicit: it documents the state transition on first success, byte-identical idempotent regeneration, that only a valid file is returned, and that transmitted is always false with two distinct reasons. This goes well beyond a typical one-line tool description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences, but each adds a distinct fact: primary action, idempotence/immutability, validity guarantee, and never-transmit routing. It is front-loaded with the main behavior and uses semicolons to keep related constraints together rather than padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and a domain-specific EBICS workflow, the description covers the critical behaviors and routing. The main gap is parameter semantics and a fuller statement of the success response beyond transmitted:false and a valid file; still, most operational decisions an agent must make are addressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only bare names for workspaceId, batchId, and idempotencyKey, and 0% of the parameters have descriptions, so the description needed to clarify them. It only refers generically to 'the batch' and does not define what each parameter must be, what idempotencyKey is used for, or any format constraints. An agent must rely on parameter names, which is insufficient for a required 3-parameter call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a concrete verb and resource: it builds or rebuilds a batch's pain.001.001.09 XML, validates it, and moves the batch to generated. It also separates itself from transmission by explicitly stating it never transmits, so it is not confused with payment_batch_transmit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the agent when transmission is out of scope and routes to the sibling: if an A33 EBICS channel exists, transmission belongs to payment_batch_transmit; otherwise no_channel and the file-download path applies. It also signals that regenerating an already-generated batch is safe and byte-identical, so the tool is appropriate for rebuilds, not just first-time generation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_agent_dialARead-only
Read the raw approval-dial levels for every governed capability (post, issue, send, dun, pay, vat-file, customize, plugin-install, close-period, go-live). Each defaults to ask when never set. Owner-facing view of what the agent may auto-execute versus draft.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description aligns with that. It adds useful behavioral detail beyond the annotation: each dial defaults to 'ask' when never set, which informs the agent about unset values. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no fluff. The enumeration of capabilities is long but directly informative, and the primary action is front-loaded. It is appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, default behavior, and audience. However, there is no output schema, and the description does not explain the return format, what 'raw' means, or the exact values of the dial levels. For a simple read tool it is acceptable, but it falls short of fully equipping the agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one required parameter (workspaceId) with 0% schema description coverage, yet the description never mentions or explains the parameter. The workspace scope is implied by context, but the description does nothing to compensate for the missing schema documentation, leaving the agent to infer the parameter's role.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Read') and a specific resource ('raw approval-dial levels'), lists the governed capabilities explicitly, and frames it as an owner-facing read view. This clearly distinguishes it from the mutating sibling set_agent_dial and from list_drafted_actions, even though siblings are not named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear context for when to use the tool: 'Owner-facing view of what the agent may auto-execute versus draft' signals it is for inspecting permissions, not changing them. It does not explicitly name alternatives or exclusions, but the read/write contrast with set_agent_dial is strongly implied by the verb.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_agent_sessionARead-only
Read one agent session in full: its turns in order, each with its recorded calls (verb, arguments, execute/draft/deny decision, dial capability, outcome, duration, and the created object where one exists). A foreign sessionId reads as not_found, never as another tenant’s transcript.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses a critical security behavior: 'A foreign sessionId reads as not_found, never as another tenant's transcript.' It also details the response structure (turns, calls, decisions, outcomes), providing behavioral context that annotations do not cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no fluff. The primary purpose is front-loaded, and the detailed field list is compactly integrated into the first sentence. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read tool with no output schema, the description covers the return contents and an important edge case. It does not mention potential large payloads, pagination, or explicitly contrast with sibling tools, but the provided information is sufficient for an agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning for sessionId by explaining the foreign-session behavior, but it does not explain workspaceId or the additionalProperties=true allowance. The parameter meanings are only partially clarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read'), a resource ('one agent session'), and the full scope ('in full'), followed by an enumeration of the session contents. This clearly distinguishes it from list_agent_sessions (which lists sessions) and get_agent_dial (which retrieves a specific dial), so an agent can tell them apart without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage when a complete session transcript is needed, but it does not explicitly name alternatives or state when not to use this tool. The presence of siblings like get_agent_dial and list_agent_sessions makes the differentiation inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_aging_bucket_configARead-only
The day thresholds the aging buckets cut at, and whether this workspace chose them or is on the shipped default of 30/60/90 days. Also returns the bucket keys those boundaries produce. Reads only.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it read-only, and the description consistently says 'Reads only' with no contradiction. It adds useful behavioral context by explaining what the returned thresholds represent, including the shipped default of 30/60/90 and whether the workspace chose custom values.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, no filler, with the core output described first and the read-only caveat last. Efficient for what it covers, though 'Reads only' slightly duplicates the readOnlyHint annotation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only tool with no output schema, the description gives enough about the returned content to invoke it correctly. It could add an explicit note about what workspaceId means, but the tool is simple enough that this is not a major gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one required workspaceId parameter with no description, and the tool description never names it explicitly. 'This workspace' implies the result is scoped by workspaceId, which partially compensates for the 0% schema coverage, but the parameter itself is left implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states what the tool returns: aging bucket day thresholds, default/custom origin, and bucket keys. The phrase 'Reads only' distinguishes it from the sibling set_aging_bucket_config without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is the tool to read aging bucket configuration, and 'Reads only' hints it is not for changing it. However, it never explicitly says 'use set_aging_bucket_config to modify these settings', leaving the contrast with its sibling mostly implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_api_catalogARead-only
Return an OpenAPI 3.x document of every MCP tool and its REST twin (route, input shape, read/write, required capability), generated from the live registry. Describes the software contract, not workspace data.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already indicates a safe read operation, and the description adds valuable context: it is generated from the live registry and explicitly notes it describes the software contract, not workspace data. This complements the annotation without contradicting it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no fluff. The primary action and object are front-loaded, and the clarifying statement about workspace data is efficient and useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only tool, the description fully explains what it returns and its scope. It does not need to describe output schema details since an OpenAPI document is self-describing, and the read-only nature is covered by annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the description does not need to explain parameter meaning. The schema coverage is trivially 100% and the baseline for zero parameters is 4, which is appropriate here.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Return') and a specific resource ('an OpenAPI 3.x document of every MCP tool and its REST twin'), and it clarifies that it describes the software contract rather than workspace data. This distinguishes it from data-retrieval tools and makes its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide any guidance on when to use this tool versus alternatives, such as the sibling 'runtime_catalog' or other introspection tools. There are no explicit conditions, exclusions, or comparisons, leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_audit_logCRead-only
Read the tamper-evident audit trail and its chain-verification status.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| entityKind | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already signals safe read behavior, so the description is not burdened with basic safety disclosure. It adds useful context around the tamper-evident nature and chain-verification status, but it doesn't explain what chain verification entails, whether verification is computed on demand, or what the response contains. No contradiction with the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The single sentence is tight, front-loaded and grammatically efficient, with no wasted words. It states the core action and the distinguishing resource in one readable line.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With four parameters, no output schema, and zero schema-level parameter descriptions, the description is too thin. It does not clarify date range semantics, allowed entityKind values, or what the chain-verification status returns, so the agent is left without enough context to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema description coverage is 0% and none of the four parameters (workspaceId, to, from, entityKind) are explained in the description. The description adds no meaning beyond the schema; the agent must guess the semantics of date-range and entityKind filters and why workspaceId is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb (Read) and a well-defined resource (tamper-evident audit trail and chain-verification status), making the tool's domain clear. It doesn't explicitly contrast with siblings like get_diagnostics or review_status, but the resource is distinctive enough to separate it from peers.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus alternatives such as review_status, runtime_status, or get_diagnostics. The description states what the tool does but leaves the selection criteria entirely to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_automation_ruleARead-only
Read one automation rule: its trigger, its condition, the verb it calls and the input template it calls it with, who authored it (and therefore whose capabilities its firings carry), and when it last fired.
| Name | Required | Description | Default |
|---|---|---|---|
| ruleId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description need not restate safety. It adds valuable context by specifying the returned data and its significance (e.g., author implies whose capabilities firings carry). It does not contradict annotations and provides useful behavioral detail beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence that front-loads the core action 'Read one automation rule' and then lists specifics. It is appropriately sized, though slightly long; every clause adds value and there is no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with no output schema, the description lists the returned fields comprehensively, including the author and last-fired timestamps. It does not cover potential error cases or null handling, but given the read-only annotation and straightforward nature, this is largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the parameters. It does not mention workspaceId or ruleId at all, leaving their purpose implicit. While they are simple string IDs, the description fails to compensate for the missing schema documentation, making it inadequate for parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Read' and the resource 'one automation rule', and enumerates exactly what is returned (trigger, condition, verb, input template, author, last fired). This distinguishes it from sibling tools like list_automation_rules (which lists many) and get_automation_run (which reads runs), so an agent can select it unambiguously.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is for reading a single rule via a ruleId, but it does not explicitly contrast with list_automation_rules or get_automation_run, nor mention when to prefer it. No exclusions or alternative conditions are given, leaving the 'when' to inference from the name and parameters.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_automation_runARead-only
Read one firing in full, including the resolved action input that was really sent rather than the template it came from.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description aligns by saying 'Read,' so there is no contradiction. The description adds behavior beyond the annotation by explaining that the return includes the resolved action input actually sent (not the template), which shapes the agent's expectation of response content. It doesn't discuss pagination, permissions, or what 'in full' includes, but for a simple read operation the behavior is adequately disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, no filler, with the most valuable differentiator (resolved input vs template) placed up front. It earns its place and is immediately parsable by an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter read operation with a readOnly annotation and no output schema, the description is mostly complete: it states the object (one firing), the level of detail ('in full'), and the key nuance (resolved action input vs template). The only small gaps are not explicitly naming what runId/workspaceId refer to and not mentioning any error conditions (e.g., run not found), but given the tool's low complexity, this is a minor shortfall.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the burden, but it only explains the semantics of runId implicitly ('one firing') and workspaceId not at all. The schema provides only provides parameter names with no descriptions. The description adds meaningful context about the resource being read but doesn't explicitly map runId to a firing. Since the schema is sparse and the description adds some clarity about the object being fetched, a 4 is justified for adding value beyond the raw schema, though it could have explicitly defined both parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Read') and resource ('one firing'), and includes a distinctive qualifier: 'in full, including the resolved action input that was really sent rather than the template it came from.' This clearly distinguishes it from sibling list/retry tools like list_automation_runs and retry_automation_run. Even without a title, the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for reading a single execution record and highlights what makes it unique (resolved inputs vs templates), which tells the agent when this tool adds value over a list tool. However, it never explicitly states when to use this versus siblings like list_automation_runs or retry_automation_run, nor does it mention any preconditions (e.g., run must exist, requires workspace context).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_bank_accountARead-only
Read one Bankkonto, including whether its IBAN is a QR-IBAN and therefore receive-only (a QR-IBAN may only be credited, never used as a debit account).
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| bankAccountId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint=true already establishes that this is a safe read operation. The description adds meaningful behavioral context by explaining that a QR-IBAN is receive-only and may only be credited, never used as a debit account, which is useful for downstream decisions. It does not contradict the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that states the core purpose first and then adds the key QR-IBAN nuance without any fluff. Every phrase earns its place, and the structure makes the tool's behavior immediately clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple get-by-two-IDs tool with a readOnly annotation, the description covers the essential purpose and the distinctive QR-IBAN implication. However, it provides no guidance on the required parameters and no indication of the return structure (other than the QR-IBAN flag), leaving some ambiguity for an agent that has no schema descriptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the undocumented workspaceId and bankAccountId parameters. It does not mention these parameters or explain what values they expect, leaving the agent to infer their purpose from the names alone. This is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Read') and resource ('one Bankkonto') and adds a distinguishing detail: it returns whether the IBAN is a QR-IBAN and therefore receive-only. This clearly differentiates it from list_bank_accounts, create_bank_account, update_bank_account, and archive_bank_account.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Read one Bankkonto' implies this tool is for retrieving a single bank account rather than listing or mutating accounts, but it does not explicitly name alternatives or state when not to use it. The QR-IBAN context hints at downstream constraints, but there is no direct exclusion like 'for all accounts, use list_bank_accounts'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_captureARead-only
One capture with all its fields (live plus superseded history, each carrying its provenance and confidence) and the E00 document reference to fetch the original via documents_get_content. A31 never serves bytes itself.
| Name | Required | Description | Default |
|---|---|---|---|
| captureId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses important behavioral traits: the response includes full history with provenance and confidence, and it explicitly states the tool never serves bytes, pointing to documents_get_content for that. This gives the agent a clear understanding of what the tool returns and its limitation, which is valuable context not captured by annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, efficiently front-loaded with the core function and includes the key behavioral caveat (no bytes) and a pointer to the alternative. Every sentence adds value, with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no output schema, the description provides a good overview of the response content (all fields, history, provenance, confidence, E00 document reference) and points to documents_get_content for the original bytes. It covers the essential aspects an agent needs to understand what the tool returns and how to use it, though it could be more explicit about the exact structure of the capture object.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides two required parameters (workspaceId, captureId) with no descriptions (0% coverage). The description does not elaborate on these parameters, leaving their meaning and format ambiguous. While captureId is implied by the tool's purpose, workspaceId is not explained, and the description fails to compensate for the lack of schema descriptions, making it difficult for an agent to know how to properly populate them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns a single capture with all its fields including live and superseded history, each with provenance and confidence. It explicitly distinguishes from sibling tools like capture_document, capture_extract, capture_commit, and capture_discard by focusing on retrieval, and it also differentiates from documents_get_content by noting that A31 never serves bytes itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: to retrieve a specific capture by ID (captureId) within a workspace (workspaceId). It also provides a critical alternative: to fetch the actual document bytes, use documents_get_content, since this tool does not serve bytes. However, it does not explicitly state when not to use this tool versus list_captures or other retrieval tools, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_company_profileBRead-only
Read the workspace company & fiscal profile.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is consistent with the readOnlyHint annotation and adds that the tool covers both company and fiscal profile data. However, no extra behavioral context is provided, such as what happens if the workspace does not exist or whether partial profile data can be returned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short sentence, front-loaded with the action and resource. Every word earns its place, and there is no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only tool this is minimally adequate, but it lacks details about the returned profile structure, any conditional behavior, or how it relates to the many workspace-related sibling tools. The absence of an output schema makes a little more description desirable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the workspaceId parameter beyond loosely implying scope through the word 'workspace'. It does not compensate for the missing schema documentation, though the single parameter is relatively self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Read' and names the resource: the workspace company and fiscal profile. This clearly distinguishes it from related sibling tools like update_company_profile or set_fiscal_config, which imply mutations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided about when to use this tool versus alternatives such as get_workspace, update_company_profile, or set_fiscal_config. The read-only nature is implied by 'Read', but no explicit conditions or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_conceptARead-only
Read one Begriff in one locale: the authored body (THE WORDING OF RECORD for explaining this term: cite it rather than paraphrasing), its statutory citations as a structured list, related terms, and an explicit not-implemented flag where TILL does not cover the subject. Unknown keys return a structured not_found naming the nearest keys; nothing is ever generated. Workspace-free: the corpus is identical in every workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| locale | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, but the description goes further by stating 'nothing is ever generated' – a strong behavioral guarantee. It also discloses how unknown keys are handled (structured not_found with nearest keys) and that the tool is workspace-free (identical corpus across workspaces). These are significant behavioral details beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately long but every sentence adds value: it lists return components, explains unknown-key behavior, and notes workspace invariance. It is front-loaded with the core purpose and avoids fluff. It could be slightly more concise, but it is well-structured and not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description adequately explains what the tool returns (authored body, citations, related terms, flag) and handles edge cases (unknown keys). It also clarifies the workspace-free nature. It does not specify the exact structure of the response (e.g., field names) but that may be inferred. The main missing piece is parameter format guidance, which was already penalized in parameter semantics. Overall, it is reasonably complete for an agent to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the parameters 'key' and 'locale'. The description mentions 'one Begriff in one locale' and 'key' implicitly, but it does not explain what constitutes a key or the valid values/format for locale. The agent is left guessing about how to specify the locale or what a key looks like. This is a clear gap given zero schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Read one Begriff in one locale' and enumerates exactly what is returned (authored body, statutory citations, related terms, not-implemented flag). It also distinguishes itself from list_concepts by focusing on a single term, though not naming the sibling explicitly. The verb 'read' plus the resource 'Begriff' makes the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this is for retrieving a single concept's details, and it explicitly advises citing the authored body rather than paraphrasing, which guides the agent's use of the result. It does not name alternative tools or provide exclusion criteria (e.g., when to use list_concepts instead), but the scope is clear enough that an agent can infer it is for single-lookup use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_contactCRead-only
Read a contact.
| Name | Required | Description | Default |
|---|---|---|---|
| contactId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes the read-only nature, and the description merely restates that without adding new behavioral details. There is no mention of what happens for missing contacts, response shape, or any side effects, but it does not contradict the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely short and front-loaded, but it is under-specified rather than appropriately concise. It avoids fluff, yet it omits essential usage and parameter context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, no parameter descriptions, and no usage guidance, this description is incomplete. It gives the agent only the bare action, not enough to confidently invoke the tool in varied scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, and the description does not explain the purpose or format of workspaceId or contactId. An agent is left to infer that these are identifiers, with no guidance on their relationship or expected values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (read) and the resource (contact), and the required contactId and workspaceId parameters make it evident this is a single-record fetch rather than a list operation. It does not explicitly distinguish itself from list_contacts, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as list_contacts, search_global, or contacts_timeline. It provides no context about typical use cases, prerequisites, or sibling tool routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_diagnosticsBRead-only
Read the error-recording preference and every recorded entry: the whole of what TILL has kept about this person, readable at any time with no request to make.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description reinforces this with 'readable at any time with no request to make.' It adds context about the scope ('every recorded entry' and 'the whole of what TILL has kept about this person'), but doesn't disclose details like return format, pagination, or data volume. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that front-loads the core action ('Read the error-recording preference and every recorded entry') and adds a clarifying scope phrase. It is efficient with no filler, though the phrasing 'the whole of what TILL has kept about this person' is slightly informal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with one parameter and readOnlyHint annotation, the description covers the basic purpose and scope. However, it leaves the workspaceId parameter unexplained and doesn't describe the output shape, which an agent would need to interpret the result correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the workspaceId parameter at all. The description mentions 'this person' but doesn't clarify how workspaceId relates to the person or what value to pass. With a single required parameter and zero schema coverage, the description should have compensated but didn't.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads the error-recording preference and all recorded entries, with a specific verb ('Read') and resource ('the whole of what TILL has kept about this person'). It distinguishes itself from set_diagnostics and clear_diagnostics by emphasizing read-only access, though it doesn't explicitly name those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is the read counterpart to set_diagnostics/clear_diagnostics and notes it is 'readable at any time with no request to make,' which gives some usage context. However, it doesn't explicitly state when to use this tool versus alternatives or mention any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_documentARead-only
Read a document with its positions and status trail. For an invoice, include 'qr' and/or 'pdf' to attach the Swiss QR-bill payload and the rendered PDF (A11, D14).
| Name | Required | Description | Default |
|---|---|---|---|
| include | No | ||
| documentId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint=true already covers the read-only nature. The description adds behavioral context by specifying the response includes positions and status trail, and the optional attachments for invoices (QR-bill payload and PDF). This goes beyond the annotation by describing the returned data and conditional behavior. No contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no wasted words. The primary purpose is front-loaded, and the conditional include information is clearly stated in a compact manner. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read tool with three parameters and no output schema, the description is fairly complete. It explains the optional includes and hints at the response content (positions and status trail). It does not describe the full response structure or any pagination or error behavior, but given the read-only nature and the clarity of the parameters, this is acceptable. A small gap is the lack of description for the required IDs, but they are obvious.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the 'include' parameter well (values 'qr' and 'pdf' and their purpose). However, it does not describe 'workspaceId' and 'documentId', though these are self-explanatory from their names. The description adds value for one param but leaves the others undocumented; this partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read'), a specific resource ('a document'), and explicitly mentions what it returns ('positions and status trail'). It also differentiates from other get tools by highlighting the optional invoice-specific includes ('qr' and 'pdf'). This clearly distinguishes it from siblings like list_documents or get_document_template.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool (to read a document with its positions and status trail) but does not explicitly state when not to use it or mention alternatives. It gives context for the include parameter for invoices, but there is no direct comparison to other document-related tools. This is adequate but not explicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_document_templateBRead-only
One template with its full configuration: footer text per locale, language mode, fixed locale, optional column order, default/archived flags, and the linked logo file id.
| Name | Required | Description | Default |
|---|---|---|---|
| templateId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers the read-only nature, and the description adds useful detail about what the returned configuration includes. However, it does not disclose any other behavioral aspects like what happens when the template does not exist, whether archived templates are included, or how linked logo details are resolved.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence that front-loads the core purpose and uses a colon list to enumerate the configuration fields. It is efficient, though the list is slightly dense and could be structured more clearly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does help by listing the configuration fields that will be returned. However, it does not account for parameter semantics, error behavior, or the relationship to sibling template tools, leaving moderate gaps for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides zero description coverage, leaving workspaceId and templateId with no explanation beyond their names. The description mentions template configuration but does not explain the parameters, their formats, or how they are used to identify the template.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool's subject as a single template and enumerates the exact configuration fields it returns (footer text per locale, language mode, fixed locale, column order, flags, logo file id). This differentiates it from list_document_templates, but it does not state an explicit verb like 'retrieves' or explicitly name the sibling it is not.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to call this tool versus alternatives such as list_document_templates, create_document_template, or update_document_template. The description implies it is for fetching one template's full configuration, but there are no explicit usage conditions or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dunning_configARead-only
The Mahnwesen policy: the three escalation levels with their days-overdue thresholds, the Mahngebühr per level (amount and fee-income account; a positive fee always books, and its VAT follows the chased invoice's own rates automatically per D69), the Verzugszins note settings (Art. 104 OR: 5% p.a. is the statutory default, only a HIGHER contractual rate is configurable), and whether this workspace chose the policy or is on the shipped defaults (1./2./3. Mahnung at 10/20/30 days overdue, no fee, no interest note). Reads only.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses specific behaviors: positive fees always book, VAT follows the chased invoice's rates per D69, only higher-than-statutory interest rates are configurable, and default thresholds. This adds valuable context about what the tool returns and how policy behaves.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph with many domain-specific clauses (Mahngebühr, Verzugszins, D69, Art. 104 OR) and enumerations. It is verbose and could be distilled to a clear summary like 'Returns the dunning policy configuration for the workspace.' Front-loading is weak; key purpose is buried in detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description does a good job of explaining the main components of the dunning policy: thresholds, fees, interest settings, and defaults. It covers essential aspects, though it might miss structural details like exact field types or currency units, but it is largely complete for understanding the tool's output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one parameter workspaceId with 0% description coverage. The description does not explain what workspaceId represents or how it is used (e.g., to identify the workspace whose policy is fetched). It relies on the name being self-explanatory, which is insufficient given the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns the Mahnwesen (dunning) policy, enumerating its components: escalation levels, thresholds, fees, interest note settings, and default vs. custom selection. The explicit 'Reads only' differentiates it from sibling set_dunning_config.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies read-only usage via 'Reads only' and the content focus, but does not explicitly mention alternatives like set_dunning_config for modifications or when to use this getter over other dunning-related tools. The guidance is implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dunning_pdfARead-only
Render one debtor's reminder letter from an issued Mahnlauf, base64-encoded: the creditor block, the overdue invoice list with days overdue (earlier Mahngebühren itemised separately from the invoice's own open amount), the Mahngebühr, the Verzugszins note, and one Swiss QR payment part PER INVOICE (amount = that invoice's open amount + the fee AS DEMANDED AT ISSUE, reference = the invoice's own QRR/SCOR, so the payment still matches the invoice). The demand FREEZES at issue (D73): every FIGURE renders from the issue-time snapshot, so a reprint states exactly the amounts that were mailed, and a fee recovered after a period lock joins the NEXT escalation letter, never this one; the creditor and debtor blocks and the QR payload's party data render CURRENT master data. Nothing is stored. A proposed run has no letter yet.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| debtorId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description details behavior thoroughly: base64-encoded output, issue-time snapshot freezing, current master data for creditor/debtor/QR party data, no persistence, and exclusion of later-recovered fees. Nothing contradicts the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause adds materially: output contents, snapshot semantics, payment-part behavior, storage behavior, and the proposed-run constraint. It is front-loaded with the core function before the detailed rules.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only render tool with no output schema, the description fully covers what the returned PDF contains and the temporal rules for fees and master data. It also flags the no-letter case for proposed runs, leaving little ambiguity for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description compensates by associating runId with an issued Mahnlauf and debtorId with the specific debtor whose letter is rendered. workspaceId remains only implicit in the surrounding domain language.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Render') and resource ('one debtor's reminder letter from an issued Mahnlauf') and enumerates the exact contents of the PDF. It clearly distinguishes the tool from siblings like get_dunning_run by clarifying it produces the rendered letter rather than metadata.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the trigger condition explicit: an issued Mahnlauf, and explicitly excludes proposed runs with 'A proposed run has no letter yet.' It does not name alternative tools, but it gives clear context for when this render operation applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_dunning_runARead-only
Read one Mahnlauf in full: every item (invoice, debtor, level, open amount, Mahngebühr, Verzugszins note, days overdue) and the per-debtor letter groups with their send outcomes. Reads only.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers the read-only safety profile, and the description adds 'Reads only.' along with a description of the returned content. It doesn't disclose potential large result sizes, pagination, or error behavior, but it does add useful content-level detail beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficient and front-loaded with the core action, then lists the specific returned fields. The final sentence 'Reads only.' is somewhat redundant with the opening verb and the readOnlyHint annotation, but the overall length is appropriate and each part contributes to understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description gives a high-level summary of the returned items and letter groups but does not define the structure or types. It doesn't mention paging, size limits, or the meaning of German terms like Mahngebühr and Verzugszins, leaving an agent with only partial information about the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description gives no meaning to the two parameters. While runId and workspaceId are self-descriptive by name, the description does not explicitly map them to the dunning run or workspace, nor does it mention required formats or how to obtain a runId.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with the specific verb 'Read' and the resource 'one Mahnlauf' (dunning run), then enumerates exactly what is included: items with invoice, debtor, level, open amount, Mahngebühr, Verzugszins note, days overdue, plus per-debtor letter groups with send outcomes. This clearly distinguishes it from list_dunning_runs and get_dunning_pdf, which exist in the sibling set.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'in full' implies this is the detailed counterpart to listing tools, but no sibling tool is explicitly named. It provides clear context on when to use it—when a complete single-run detail is needed—but doesn't explicitly state when not to use it or mention alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_ebill_configARead-only
Read the workspace eBill configuration (the eBill-Biller-ID the owner received from a certified network partner at enrollment), or null when none is set. Enrollment happens with the partner, not in TILL.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations include readOnlyHint=true, so the description doesn't need to restate that. It adds useful context about the return value (null when none set) and the enrollment process, but doesn't disclose details like whether the ID is masked or if there are any side effects. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no waste. The core behavior and return value are front-loaded, and the enrollment clarification is a useful addition without bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with one parameter and readOnlyHint annotation, the description covers the essential behavior, return value, and a key domain fact (enrollment happens externally). It doesn't describe error cases or response format, but the output schema is absent and the tool is simple enough that this is acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden. It explains the workspaceId parameter implicitly by referring to 'workspace eBill configuration', but doesn't explicitly describe the parameter format or constraints. The single parameter is simple enough that the description adds adequate meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads the workspace eBill configuration and returns the eBill-Biller-ID or null. It specifies the resource (workspace eBill configuration) and the exact return value, distinguishing it from set_ebill_config and other eBill tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is a read-only retrieval tool for checking the eBill configuration, and notes that enrollment happens with the partner, not in TILL. It doesn't explicitly name alternatives or exclusions, but the context makes it clear when to use it versus set_ebill_config.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_entryBRead-only
Read a journal entry and its lines. Each line reports the transaction amounts (debit, credit) under its currency, and the amounts the books hold (baseDebit, baseCredit) under baseCurrency, which is the workspace base currency.
| Name | Required | Description | Default |
|---|---|---|---|
| entryId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description supplements the readOnlyHint annotation by detailing what information is returned: each line includes transaction amounts (debit, credit) under the line's currency and base amounts under the workspace base currency. This adds meaningful context about the output structure beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the primary action, and every sentence contributes useful information about the returned data. No redundant or irrelevant content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main output elements but omits parameter explanation and any note on pagination or limits. For a simple read operation with a clear read-only annotation, it is adequate but not fully comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only parameter names and types, with 0% description coverage. The tool description does not explain the purpose or format of workspaceId and entryId, which are required parameters. This is a significant gap for an agent needing to construct a valid request.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as reading a journal entry and its lines, with a specific verb and resource. However, it does not explicitly distinguish itself from sibling tools like list_journal, though the resource type is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool instead of alternatives such as list_journal or get_document. There is no mention of preconditions, related tools, or scenarios where this is preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_exchange_rateARead-only
Resolve which rate WOULD convert a currency into the workspace base currency on a date, and report the validity date, source and method of the rate that governs. Answers needs_fx_rate when none is admissible, rather than inventing one.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | ||
| currency | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already signals no mutation, and the description adds valuable behavioral context: it reports the validity date, source, and method of the governing rate, and it explicitly answers needs_fx_rate rather than inventing a rate. This is meaningful beyond the annotation and contains no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the first sentence states the core behavior and the output context, and the second clearly communicates the fallback behavior. Every sentence earns its place without repeating schema field names.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description usefully indicates what the tool returns: the rate plus its validity date, source, and method, or needs_fx_rate. It could be slightly more explicit about whether the numeric rate value itself is returned, but overall it is largely sufficient for a small read-only tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the semantic load and does clarify that currency is the source currency, workspaceId identifies the workspace base currency, and date is the relevant conversion date. However, it does not clarify date optionality/defaults or expected format details, leaving some ambiguity for invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific operation: resolving which rate would convert a currency into the workspace base currency on a date. It also distinguishes the tool from rate CRUD siblings like record_exchange_rate and list_exchange_rates by framing it as a resolution/calculation rather than a create/list operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The use case is clearly stated: determine the governing rate for a currency conversion on a specified date, with a fallback of needs_fx_rate when no rate is admissible. It does not explicitly name alternative tools or exclusion conditions, but the context is strong enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_fx_methodARead-only
Report which MWSTV Art. 45 conversion basis governs a Steuerperiode, whether that period is already settled by a posted foreign-currency entry, the earliest period whose basis may still be chosen, and the full election history. Ask by taxPeriod (YYYY) or by any date inside it.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | ||
| taxPeriod | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already set readOnlyHint=true, so the description only needs to add context beyond that. It does, by disclosing what kind of report is returned: governing basis, settlement state, earliest electable period, and full election history. It does not contradict the annotations and adds meaningful behavioral context, though it does not discuss authorization or response semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The dense first sentence front-loads the full scope of the report, and the second sentence gives the query instructions in a scannable format with code-styled parameter names.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is strong for a read-only query: it lists all major outputs and the supported lookup modes, which is especially important given there is no output schema to describe returns. It does not specify exact response formatting or whether at least one of `taxPeriod`/`date` must be supplied, but for a read report tool the description is largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates meaningfully: it defines `taxPeriod` as YYYY and explains that `date` can be any date inside the period. It leaves `workspaceId` undocumented, but that parameter is standard and the name is self-explanatory. The 'Ask by... or by...' phrasing also communicates the two query alternatives.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb ('Report') and a specific resource: the MWSTV Art. 45 conversion basis for a Steuerperiode. It then enumerates three additional facts (settlement status, earliest choosable period, election history), making the tool's purpose unmistakable and distinct from siblings like set_fx_method or describe_rate_feed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the caller how to query: 'Ask by `taxPeriod` (YYYY) or by any `date` inside it.' That is clear usage guidance. It does not name alternatives or state when not to use this tool, so it stops short of the explicit when/when-not standard.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_itemDRead-only
Read an item.
| Name | Required | Description | Default |
|---|---|---|---|
| itemId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The only behavioral signal is the readOnlyHint annotation, which the description implicitly repeats without adding extra context. It does not disclose return format, scoping constraints, or error behavior. Beyond the annotation, no new information is provided.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
At three words, the description is minimal but under-specified. It fails to convey essential context beyond the name, so this is under-specification rather than concise effectiveness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema, and the description does not explain what an 'item' is in this domain, how the response looks, or any special behavior. Given the array of sibling tools, this minimal description is insufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the meaning or format of workspaceId or itemId. With no descriptions in the schema and no compensation in the description, the agent must infer the semantics from parameter names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Read an item' restates the tool name with synonyms ('read' = 'get', 'item' = 'item') and does not specify what kind of item is being retrieved. It does not distinguish from sibling getters like get_document or get_contact. This is essentially a tautology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternative getter tools. It does not mention contexts, prerequisites, or exclusions. An agent has no basis for choosing this over get_document or get_contact.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_move_stateARead-only
Read the move-checklist resume point for this workspace: the direction the books are moving (local_to_selfhost | local_to_managed | selfhost_to_managed | selfhost_to_local | managed_to_local | managed_to_selfhost), the five checklist steps each with the instant it was checked off (or null), startedAt and completedAt; or move:null when no move was started. Pure bookkeeping: nothing in the engine gates on this row. A COMPLETED move keeps its record, which is what the stale-writable notice (a moved but unarchived ledger) is derived from.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, consistent with the description's 'Read' and 'Pure bookkeeping.' The description adds valuable nuance beyond the annotation: nothing gates on this row, a completed move retains its record, and the stale-writable notice is derived from this data. No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every clause earns its place: the action and resource are front-loaded, the returned fields are enumerated precisely, the null case is covered, and the two contextual notes (no gating, stale-writable derivation) are high-value. No filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read with readOnlyHint, the description is nearly complete: it specifies the return contents, the null case, and the lifecycle nuance of a completed move. It does not describe the JSON response envelope or error behavior, but with no output schema the field enumeration is especially useful. The only real gap is explicit routing among sibling move/checklist tools, already noted under usage guidelines.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for workspaceId, and the description only refers to 'this workspace,' which confirms the parameter identifies the workspace but adds no format, constraint, or additional context. For a single self-evident required parameter this is minimally adequate, though the description does not compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Read') and resource ('move-checklist resume point'), then enumerates the exact data fields returned: direction enum, five checklist steps with timestamps, startedAt/completedAt, or move:null. The 'pure bookkeeping' note and the stale-writable derivation clearly separate it from mutation and migration siblings, so an agent can distinguish it from advance_move_step or checklist_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: it is the read-only inspection point for a move, and 'nothing in the engine gates on this row' implies it is safe for diagnostic use. It also explains that a completed move keeps its record and that the stale-writable notice is derived from it, which tells the agent when this tool is relevant. However, it does not explicitly name an alternative or state when-not-to-use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_onboarding_progressARead-only
Read the first-run wizard resume point for this workspace ({path, step, completedAt} or null when the wizard never ran), plus the workspace kind (demo|sandbox|live) so a client can show the demo banner off the same call. Pure bookkeeping: nothing in the engine gates on this row.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description adds 'Pure bookkeeping: nothing in the engine gates on this row' and clarifies the null return case, providing useful behavioral context. It does not contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler, front-loading the core function and return shape, then adding a purpose note. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one required parameter and no output schema, the description fully describes the return structure and the workspace kind. It omits error handling or permission notes, but these are not critical for a simple read operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for workspaceId, and the description only refers to 'this workspace' without explicitly explaining that workspaceId identifies the target workspace. For a single parameter, the description provides minimal semantic value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads the first-run wizard resume point and workspace kind, specifying the exact return shape ({path, step, completedAt} or null). It uses a specific verb (Read) with a concrete resource, and the added context about the demo banner distinguishes its purpose from other workspace-related reads.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage for showing the demo banner and explicitly notes 'Pure bookkeeping' to signal it's a safe read. However, it does not name alternatives (e.g., get_workspace) or state when not to use it, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_opening_balancesARead-only
Read the opening position for a fiscal year: the seeded entry when there is one, or the balance sheet carried from the prior year's close (source=carried_forward, editable=false), or none. Returns each account with its number, name, type and side in Rappen, plus the totals and the reference (the Inventar or Beleg the position is traced to, null when none was given). A position that was reversed reads as source=none, because the ledger no longer holds it. A carried position is derived from the close and is not re-keyable: A03 owns the year-end result posting. If a carried position does not tie out it is REFUSED with carried_position_unbalanced rather than handed on, naming the signed difference and unclosedYears: an unclosed prior year leaves its result on Erfolgsrechnung accounts that a balance sheet does not carry, and every report above this one would be built on the gap.
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Even though readOnlyHint=true is already in the annotations, the description adds substantial behavioral context: carried positions are derived from the close and not re-keyable, A03 owns the year-end posting, reversed positions read as source=none, and unbalanced carried positions are refused with a specific error. It also explains the reasoning behind the refusal. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence adds meaningful detail about behavior, return values, and edge cases. It is front-loaded with the core purpose and progressively elaborates. While it is long, the complexity of the tool justifies the length; there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully covers the return format (accounts with number, name, type, side in Rappen, totals, reference) and the various source states. It also explains the refusal scenario and its cause, which is critical for an agent to interpret errors. Nothing needed for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage (0%), so the description must compensate for parameter documentation. It mentions 'a fiscal year', implying the 'year' parameter, but it never explicitly describes 'workspaceId' or clarifies the expected format (e.g., YYYY, string). The description adds minimal parameter semantics beyond what the property names already suggest.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read') and a precise resource ('the opening position for a fiscal year'), and distinguishes the three possible sources (seeded entry, carried forward, none). It clearly differentiates from write-oriented siblings like set_opening_balances, import_opening_balances, and preview_opening_import.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly conveys that this is a read-only query for opening balances, and implicitly contrasts with the write tools (set_opening_balances, import_opening_balances) present in the sibling list. However, it does not explicitly name an alternative or state a 'when not to use' condition, so it stops short of full guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_paymentARead-only
Read one payment: its allocations by document number, its Guthaben, its journal entry, and its reversal when it has one.
| Name | Required | Description | Default |
|---|---|---|---|
| paymentId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description's 'Read' verb is consistent with that, so no contradiction. The description adds useful context about what the call surfaces (allocations, Guthaben, journal entry, conditional reversal) beyond the annotation. With the read-only safety profile already covered by annotations, the bar is lower, and the description meets it without going further into auth or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the verb and resource and then lists the returned components with no filler. The conditional phrasing 'and its reversal when it has one' adds precision without waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with two straightforward parameters and no output schema, the description covers the essential ground: what it reads, for which entity, and what comes back. The main gaps are the undefined domain term 'Guthaben', no mention of missing-payment behavior, and no sibling routing, but these are minor for the tool's modest complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it adds little: 'Read one payment' implicitly ties paymentId to the payment being fetched, and workspaceId is left entirely unexplained. The parameter names themselves are conventional and self-explanatory among sibling tools, which keeps this at a baseline 3 rather than lower.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Read one payment') and enumerates the response composition (allocations by document number, Guthaben, journal entry, reversal when present). The singular 'one payment' clearly distinguishes it from list_payments and preview_payment among the siblings, so an agent can route correctly without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Read one payment' phrasing conveys an implied usage context: use this when you need a single payment's full details rather than a list. However, it never explicitly names alternatives like list_payments or preview_payment, nor does it state when not to use it, so the routing guidance is inferred rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_payment_batchARead-only
Read one payment batch in full: its debtor account, execution date, status, control sum and every item with its vendor, amount, full and masked creditor IBAN (the destination a pre-upload review verifies) and settlement.
| Name | Required | Description | Default |
|---|---|---|---|
| batchId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation declares readOnlyHint=true, which is consistent. The description adds that the batch is read 'in full' and clarifies the IBANs are the destination for pre-upload review, which is extra context. However, it doesn't disclose potential return size or that additionalProperties are allowed. With the annotation covering safety, the description adds some value but not exhaustive behavioral detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence that front-loads the core action and then lists key returned fields. It is efficient with no fluff, but the parenthetical about pre-upload review inserts a purpose clause that, while useful, interrupts the flow. Still, it's concise and clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no output schema, the description lists the main fields returned, which is helpful. It does not mention pagination or the exact format of items, but the tool likely fits a standard batch structure. The sibling list includes related tools, and the description is sufficient for an agent to select and call it correctly, though it could note that it's read-only (already covered by annotation).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters. It implies that workspaceId and batchId identify the batch, but doesn't explicitly state their types or formats. However, the description clarifies what the tool returns, which indirectly helps with expected inputs. Given 0% coverage, it partially compensates, but could add more explicit parameter guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads one payment batch in full, enumerating specific fields (debtor account, execution date, status, control sum, items with vendor, amount, IBANs and settlement). This verb+resource combination distinguishes it from siblings like list_payment_batches, create_payment_batch, and mark_batch_paid.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
While the read-only nature is implied and the scope is clear, there is no explicit statement of when to use this tool versus alternatives like list_payment_batches or generate_pain001. The description mentions a specific use case (pre-upload review), but does not say 'use this when...' or 'use list_payment_batches for summaries'. The usage is clear but not contrasted with siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_pluginARead-only
Lies eine Erweiterung (US-G02.2/3/4): one plugin manifest with its granted permission set and the precise list of capability registrations it currently holds (empty when disabled or incompatible).
| Name | Required | Description | Default |
|---|---|---|---|
| pluginId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations provide readOnlyHint=true, so the description does not need to restate that. The description usefully adds that capability registrations are empty when disabled or incompatible, which explains a behavioral nuance. However, it doesn't disclose return format, required scopes, or what happens if pluginId does not exist. Overall, it adds meaningful context beyond annotations without contradicting.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one sentence with only relevant information, front-loading the verb and resource. It adds the business rule about empty capability registrations without waste. Every part is meaningful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With readOnlyHint annotation, no output schema, and 2 required parameters, the description covers the main purpose and edge case of disabled/incompatible plugins. It remains incomplete about response shape or error behavior, but for a simple read operation the essentials are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description names the resource and clarifies that it reads a specific plugin manifest, implying pluginId is the plugin identifier and workspaceId scopes the workspace. It does not, however, provide detailed meaning or format for either parameter, such as where to find workspaceId or pluginId. The baseline of 3 applies as the description adds limited meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Lies' (reads) with a specific resource 'Erweiterung' (plugin) and specifies exactly what is read: one plugin manifest, granted permission set, and capability registrations (empty when disabled or incompatible). This clearly distinguishes it from siblings like list_plugins, install_plugin, enable_plugin, disable_plugin, and search_plugin_registry.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description specifies what the tool does but does not explicitly say when to use this over alternatives such as list_plugins or get_plugin_registry_entry. Its scope is implied by naming the resource and the details it returns, but there is no explicit when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_plugin_registry_entryBRead-only
Lies einen Verzeichnis-Eintrag (US-G02.5): resolve one registry reference to its full entry (its requested scopes drive the install-review dialog). needs_registry when no registry is configured, registry_unreachable when a configured one does not respond (P9).
| Name | Required | Description | Default |
|---|---|---|---|
| registryRef | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds specific error conditions (needs_registry, registry_unreachable) which is useful behavioral context. It does not contradict the annotation and provides additional detail beyond the basic read-only nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately concise but includes extraneous elements like 'US-G02.5' and 'P9' which are not self-explanatory and may confuse the agent. The core information is present but could be streamlined.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and 0% parameter coverage, the description is incomplete. It does not describe the return value (the full entry structure) or clarify the meaning of the parameters. The error conditions are helpful but do not compensate for the lack of parameter and return value documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the parameters. It mentions 'registry reference' and 'workspaceId' but does not describe their meaning, format, or constraints. The description refers to a 'registry reference' without clarifying what that is, leaving the agent to infer from context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the core action: 'resolve one registry reference to its full entry' and mentions that requested scopes drive the install-review dialog. It is specific enough to distinguish from a search tool, though it does not name the sibling directly. The German opening adds context but is not strictly necessary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is for resolving a single reference, and mentions error conditions (needs_registry, registry_unreachable) but does not explicitly state when to use this tool versus alternatives like search_plugin_registry. No when-not-to-use guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_recurring_scheduleARead-only
Read one schedule in full: its snapshotted template, cadence and bounds, plus the run log (every generated period with its outcome and, where the document still exists, its number and status), which is where a generated invoice states its provenance.
| Name | Required | Description | Default |
|---|---|---|---|
| scheduleId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds meaningful behavioral context beyond that: it reveals that the schedule stores a 'snapshotted template' (not a live reference), that the run log includes every generated period with its outcome, and that document provenance is only included 'where the document still exists' – a conditional behavior an agent needs to know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, dense sentence that front-loads the core action ('Read one schedule in full') and then packs the return contents in a well-structured list. Every clause earns its place: the snapshotted template, cadence, bounds, run log, and the provenance condition. No filler or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only getter with two simple string parameters and no output schema, the description covers the essential return contents and the key conditional (document existence). It doesn't describe pagination or the shape of the run log entries, but for a single-schedule read with no output schema, the description is sufficiently complete for an agent to call it correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of explaining what the parameters mean. It does not explicitly explain workspaceId or scheduleId, but the tool name and the phrase 'one schedule' make it clear that scheduleId identifies the schedule and workspaceId scopes it. The description adds no format or constraint details beyond what the schema already shows (both are plain strings, both required). Baseline 3 is appropriate because the schema is minimal and the description doesn't compensate with extra parameter-level guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read') and a specific resource ('one schedule in full'), and enumerates exactly what is returned: the snapshotted template, cadence and bounds, and the run log with per-period outcomes and document provenance. This clearly distinguishes it from sibling tools like list_recurring_schedules (which lists schedules) and run_due_recurring (which executes them).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the read-only nature and full-detail scope explicit, which tells an agent when to use this tool: when it needs the complete schedule plus run log rather than a list entry. It does not explicitly name alternatives or state when not to use it, but the contrast with list_recurring_schedules is strongly implied by 'one schedule in full' and the run-log detail.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_sync_contractARead-only
Describe the till-sync publish contract for this workspace: the contract majors this build speaks, whether publishing is on, the current stream head and epoch, and the §I sole-writer / one-file-per-tenant posture. Pass contractVersion to negotiate: an unknown major is refused (unsupported_contract) rather than best-effort parsed.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| contractVersion | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds non-obvious behavioral context: an unknown major produces unsupported_contract rather than being leniently parsed. This is genuinely useful beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the return contract contents and ending with the negotiation caveat. No filler, no repetition of schema or annotation data.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description does a decent job naming the high-level return fields and the main failure mode. Still, it leaves gaps: workspaceId semantics, contractVersion format, and representation of stream head/epoch are unspecified, so an agent may still need to guess.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. contractVersion gets real semantics: it negotiates and can be refused. However, workspaceId is not described at all, and the expected format for contractVersion is not specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Describe') and a specific resource ('till-sync publish contract for this workspace'), and enumerates the fields returned: contract majors, publishing status, stream head/epoch, and sole-writer posture. It does not explicitly name sibling sync tools, so differentiation is clear by content but not by explicit contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives one usage instruction: pass contractVersion to negotiate, and warns that unknown majors are refused rather than best-effort parsed. However, it does not say when to choose this tool over sibling tools like sync_stream_status, sync_stream_read, or sync_publish_enable/disable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_vendor_billARead-only
Read one vendor bill: its figures and stored input-VAT trace, the journal entry it posted (and the reversal when it has one), its derived settlement status and open amount, and every payment that touched it.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| vendorBillId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description does not contradict this. It adds meaningful behavioral detail beyond the annotation by specifying exactly what data is returned (input-VAT trace, journal entry, reversal, settlement status, open amount, payments), which helps the agent predict the output without an output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that is information-dense but not overly verbose. It leads with the core action and then enumerates the specific data items, which is efficient and well-structured. No fluff or redundancy is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no output schema, the description compensates by listing all returned elements, making the tool's output predictable. It covers the essential aspects of a single-entity read operation. Minor gaps include no mention of error conditions or pagination, but these are not critical for a read of one bill by ID, and the annotations cover the safety profile.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of clarifying parameters. Although it doesn't explicitly describe workspaceId or vendorBillId, the phrase 'one vendor bill' makes it clear that vendorBillId identifies the target bill, and workspaceId follows the common pattern seen in sibling tools. The description provides minimal added meaning beyond the schema's field names, but is sufficient given the tool's straightforward nature.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (Read) and the resource (one vendor bill), then enumerates the specific data returned (figures, VAT trace, journal entry, reversal, settlement status, open amount, payments). This distinguishes it from related tools like list_vendor_bills and get_document, leaving no ambiguity about its scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for retrieving a single vendor bill's full detail, but it does not explicitly mention when to prefer this over alternatives such as list_vendor_bills (for multiple bills) or get_document (for general documents). No exclusions or alternative conditions are provided, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_workspaceARead-only
Read one workspace roster entry for the switcher header: name, legal form, base currency, fiscal year start, archived flag and creation date. Lighter than get_company_profile, which returns the full fiscal and creditor configuration.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already declares this as a read operation, and the description adds value by listing the returned fields and comparing it to get_company_profile. No contradiction; it provides context beyond the annotation without over-explaining.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose and field list, then a brief comparison. No wasted words; every clause contributes to understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read with one parameter and no output schema, the description covers the returned fields and the distinction from get_company_profile. It does not mention error conditions or edge cases, but these are not critical for a read-only roster lookup, and the annotation covers the safety profile.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description should compensate by explaining the workspaceId parameter, but it only uses the parameter name without any guidance on format, source, or validation. The name is self-explanatory, but no additional semantic context is provided, leaving the agent to infer its meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Read'), the resource ('workspace roster entry'), and enumerates the specific fields returned (name, legal form, base currency, fiscal year start, archived flag, creation date). It also distinguishes itself from get_company_profile, making its scope unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names get_company_profile as an alternative and explains that this tool is lighter, implying when to use it (roster-level info) vs. the heavier alternative. However, it does not mention other relevant siblings like list_workspaces for listing multiple entries, so the guidance is good but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gl_archive_account_historyARead-only
Aggregate one account's archived history per month, quarter or year (debit, credit, running balance), by live target account id or by the source system's own account number. Aggregates only: at 300000 entries the archive answers with periods, not rows.
| Name | Required | Description | Default |
|---|---|---|---|
| groupBy | No | ||
| accountId | No | ||
| workspaceId | Yes | ||
| sourceAccount | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, so the read-only nature is covered. The description adds useful behavioral context beyond that: it warns that this tool is aggregate-only and that at 300000 entries the archive responds with periods rather than rows, which is non-obvious and relevant for callers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, with the core behavior stated in the first sentence. The second sentence is slightly cryptic but still informative, warning about the aggregate-only behavior at scale.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description gives enough context for selection and basic invocation, including grouping granularity, identifier modes, and output concepts like debit, credit, running balance, and periods. But it omits parameter-wise specifics such as required/optional dependencies, exact groupBy values, and any behavior when multiple identifiers are provided. With no output schema, those gaps matter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the burden for parameters. It does clarify that accountId is the 'live target account id' and sourceAccount is the 'source system's own account number', and that groupBy relates to month/quarter/year granularity. However, workspaceId is not mentioned and no exact allowed values or precedence rules are given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('Aggregate'), the resource ('one account's archived history'), the grouping options ('per month, quarter or year'), and the output fields ('debit, credit, running balance'). It also distinguishes itself from row-level archive queries by emphasizing 'Aggregates only'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when aggregate archived history for a single account is needed, rather than row-level details. However, it never explicitly names alternatives such as gl_archive_query or gl_archive_periods, nor does it state when those should be preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gl_archive_importA
Import the prior system's journal into the read-only GL archive (GeBüV Art. 10: readable across the system change). Writes gl_archive rows ONLY, posts NOTHING to the live ledger, and never mixes with it: archive rows are outside the TILL Belegkette and say so. Unmapped accounts import with a null target; unbalanced source entries import flagged, never corrected. Idempotent: a re-run replaces the step's rows wholesale.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| stepId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden and does so excellently. It reveals the isolation behavior, the handling of unmapped accounts, the treatment of unbalanced entries, and the idempotency/replacement behavior. This is far beyond what the schema alone reveals.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three dense sentences with no filler. It front-loads the primary purpose, then covers isolation, edge-case behavior, and idempotency. Every sentence adds distinct value and the structure is logical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter, side-effect-heavy import tool with no output schema, the description covers the most important operational context: write scope, ledger isolation, data-quality handling, and rerun semantics. Minor gaps remain around how results or errors are surfaced and how the source journal is supplied, but these do not prevent a capable agent from understanding the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It indirectly clarifies stepId by saying a re-run 'replaces the step's rows wholesale' and idempotencyKey through the idempotent behavior, but it does not explain workspaceId or planId semantics. The parameter names are fairly conventional, but the description does not fully carry the parameter-documentation burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Import the prior system's journal into the read-only GL archive'. It further distinguishes the tool from live-ledger operations by stating it 'Writes gl_archive rows ONLY, posts NOTHING to the live ledger' and from other archive tools by explicitly framing it as an import operation, not a query or purge.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly situates when this tool should be used: for the prior system's journal during a system change, targeting the read-only GL archive. It also implies when not to use it by emphasizing that it never posts to the live ledger or mixes with it. It does not name sibling alternatives explicitly, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gl_archive_periodsBRead-only
List the archived periods with their entry counts, their OR 958f retention dates, and any purge records (completed purges AND refusals on retention grounds, each citing the statute).
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers the safe read-only nature. The description adds specifics about the returned data (entry counts, retention dates, purge records) but does not disclose any other behavioral traits like pagination, ordering, or permission requirements. This adds some value but not substantial beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence that front-loads the main action and lists the key data elements without unnecessary fluff. It is appropriately sized for a straightforward list operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without an output schema, the description partially explains the return content (entry counts, retention dates, purge records) but does not specify structure, ordering, or pagination. The workspaceId parameter is not explained, and no error or edge cases are mentioned. For a simple list tool with one parameter, this is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the single parameter workspaceId. However, the description does not mention the parameter or its meaning. The parameter name is self-explanatory to some degree, but the description adds no clarification, which is a significant gap given the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List'), a clear resource ('archived periods'), and details the exact data returned (entry counts, OR 958f retention dates, and purge records). This distinguishes it from sibling tools like gl_archive_query and gl_archive_preview, which have different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as gl_archive_query or gl_archive_preview. There is no mention of context, prerequisites, or exclusions, leaving the agent to infer usage from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gl_archive_previewARead-only
Preview a gl_history step's archive import without writing anything: per-period entry and line counts, per-account totals, the unmapped source accounts (they would import flagged, never refused) and the internally unbalanced entry count.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| stepId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation already declares readOnlyHint=true, and the description reinforces this with 'without writing anything'. It adds valuable behavioral detail beyond the annotation: unmapped source accounts 'would import flagged, never refused' and it reports the internally unbalanced entry count.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence that immediately establishes the tool's purpose and non-writing nature. The subsequent list is dense but every item adds relevant information; there is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description correctly bears the burden of describing what the preview returns; it names per-period counts, per-account totals, unmapped source accounts, and unbalanced entry count. Parameter semantics remain unexplained, but that gap is separate and was already accounted for.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description does not compensate by explaining workspaceId, planId, or stepId. The parameter names are somewhat self-explanatory, but the description adds no meaning beyond what the bare schema already exposes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Preview'), a specific resource ('a gl_history step's archive import'), and explicitly says it writes nothing. It also enumerates the concrete outputs, so an agent can clearly distinguish it from mutating/import/query tools like gl_archive_import and gl_archive_query.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this is a non-writing preview of an archive import, which implies it should be used before committing an import. It does not explicitly name alternatives or state when-not-to-use it, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gl_archive_purgeA
Purge archived periods whose OR 958f ten-year retention has expired, with a recorded reason. Deletes ARCHIVE data only (structurally incapable of touching a live journal row), keeps a purge record where the periods were, and refuses inside the retention window with the date, recording the refusal: that record is the answer a data subject receives. Requires confirmed:true.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| periodTo | Yes | ||
| confirmed | No | ||
| periodFrom | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does exceptionally well: it explicitly states that only ARCHIVE data is deleted, that live journal rows are structurally untouchable, that a purge record is kept, and that refusals inside the retention window are recorded with the date. This is exactly the behavioral disclosure a destructive tool needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the core purpose, then safety and refusal behavior. It is longer than a single sentence but every sentence earns its place; the only slight blemish is the unexplained 'OR 958f' abbreviation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive purge tool with no annotations and no output schema, this description covers the essential operational context: what gets deleted, what is kept, refusal behavior, and confirmation. It does not explain idempotency key semantics or return value, but those are minor against the strong behavioral coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning for 'confirmed' (requires true), 'reason' (recorded), and the period range (purge archived periods), but it leaves 'idempotencyKey' and date format unaddressed. Partial compensation but not enough for six parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('purge') and resource ('archived periods') with a precise condition (ten-year retention expired). This clearly differentiates the tool from gl_archive_query, gl_archive_preview, or gl_archive_import, since it is the only one that deletes archived periods.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states when to use the tool (when retention has expired), what is required (confirmed:true), and describes the refusal behavior inside the retention window. It does not explicitly name alternatives or say 'use X instead', but the context is clear enough for an agent to decide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gl_archive_queryARead-only
Query the prior-system archive by account (target or source), period, text and amount range, paginated. Returns ARCHIVE rows only, labelled with their provenance (source system, covered range): a live journal row can never appear here, and an archive row never appears in list_journal. savedViewId applies a saved view (G00).
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| page | No | ||
| text | No | ||
| account | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| balancedOnly | No | ||
| sourceAccount | No | ||
| amountMaxMinor | No | ||
| amountMinMinor | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds behavioral context beyond the readOnlyHint annotation: it returns only ARCHIVE rows, labels them with provenance (source system, covered range), and applies saved views. It also clarifies the mutual exclusivity with list_journal. These details are not present in the annotation and help the agent understand the tool's behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core query criteria and pagination, followed by behavioral context and saved view mention. Every sentence adds value, with no fluff or redundancy. It is efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the 11 parameters, no output schema, and no parameter descriptions, the description covers the main filtering dimensions and clarifies the archive-vs-live distinction. It does not mention all parameters (balancedOnly, workspaceId), nor does it describe the return format in detail (though provenance is mentioned). However, it is sufficiently complete for an agent to understand how to invoke the tool for typical archive queries.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by listing the key filter dimensions: account (target or source), period (likely from/to), text, amount range (amountMin/Max), pagination (page), and savedViewId. It does not explain every parameter (e.g., balancedOnly, workspaceId), but it provides enough semantic meaning for the main filtering capabilities. The description partially bridges the schema's lack of parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (query), the resource (prior-system archive), and the filtering criteria (account, period, text, amount range). It explicitly contrasts with list_journal, stating that archive rows never appear there, which differentiates the tool from that sibling. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: use this tool for archive rows, and use list_journal for live journal rows, by stating the mutual exclusivity. It also mentions savedViewId usage (G00). However, it does not explicitly name list_journal as the alternative tool to use for live rows, nor does it provide conditions for other archive-related siblings like gl_archive_preview or gl_archive_import. The guidance is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goods_receipt_accept_linesA
Release HELD lines (inspectionStatus pending) into stock after a quality check: exactly what the post would have done for them, at the receipt own receivedAt and under the same live open-quantity and period-lock guards. Writes the J02 movement, the trail row and the po_line.received_qty increment per accepted line, and records an append-only acceptance event with the actor. A line that is not pending is invalid_transition.
| Name | Required | Description | Default |
|---|---|---|---|
| grId | Yes | ||
| reason | No | ||
| lineIds | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses concrete side effects (J02 movement, trail row, received_qty increment, append-only acceptance event), the guard conditions, and the invalid_transition error. Since annotations are absent, this carries the burden and largely succeeds, though idempotency semantics and reversibility are not mentioned.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the primary action, then side effects and error condition. Every sentence earns its place; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Captures the tool's essence and side effects but omits expected return value (no output schema), partial-failure behavior, and idempotency details. For a complex mutation with no annotations, this is functional but incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate, but it only implicitly references lineIds and does not explain idempotencyKey, reason, workspaceId, or grId. The meaning of some parameters can be inferred, but the required idempotency key is unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Release'), resource (HELD lines in inspectionStatus pending into stock), and distinguishes from siblings by framing it as exactly what goods_receipt_post would do for these lines, plus the invalid_transition guard. This is clear enough for an agent to select it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: release after quality check, with the same guards as post. Implicitly contrasts with reject_lines by being the acceptance counterpart, but does not explicitly name alternatives or exclusion conditions. No when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goods_receipt_cancelA
Abandon a DRAFT goods receipt (nothing physical has happened yet, so nothing has to be unwound). A posted receipt cannot be cancelled: use goods_receipt_reverse, which leaves the audit trail intact. Cancelling anything other than a draft is invalid_transition.
| Name | Required | Description | Default |
|---|---|---|---|
| grId | Yes | ||
| reason | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full responsibility for behavioral disclosure. It explains the draft vs. posted distinction, that nothing physical has happened yet requiring unwinding, and that cancelling a non-draft is invalid. It does not mention effects like deletion or idempotency, but the provided context is substantial for a state-dependent operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences long, with the core action front-loaded, a clear contrast to the alternative tool in the second sentence, and a concise validity constraint in the third. Every sentence earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the key state constraint (draft only), the alternative for posted receipts, and the error condition. It doesn't mention return values or auth requirements, but the schema covers required parameters and there is no output schema. For a cancellation operation, the context is sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the four parameters. It does not mention grId, reason, workspaceId, or idempotencyKey at all)Skip? Actually, it only talks about the operation's semantics. Even if param names are self-explanatory, the description adds no value beyond the schema's field names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and object: 'Abandon a DRAFT goods receipt,' making the tool's function unambiguous. It also distinguishes the tool from goods_receipt_reverse by explicitly stating that a posted receipt cannot be cancelled. This clearly differentiates it from the sibling tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: only for draft goods receipts, and directly points to goods_receipt_reverse for posted receipts, explaining that reverse preserves the audit trail. It also states the error condition for invalid usage ('invalid_transition'), leaving no ambiguity about scope.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goods_receipt_createA
AGENT-FIRST. Open a DRAFT Wareneingang against an open purchase order, the first-class document that records what physically arrived. receivedAt (ISO date) is THE date: every stock movement this receipt ever writes is stamped with it and the period lock is checked against it here and again at post, so there is no way to back-charge a sealed year by moving a posting date. defaultLocationId is optional (a single-location workspace resolves the default warehouse location automatically). Nothing physical happens yet: no stock movement, no purchase-order quantity change. Refused when the order is not sent or received (invalid_transition), when no line has open quantity (nothing_open), or when receivedAt falls in a locked period (period_locked). Idempotent under idempotencyKey.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| poId | Yes | ||
| expectedAt | No | ||
| receivedAt | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| defaultLocationId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It excellently discloses: receivedAt is the definitive date for stock movements and period lock checks; defaultLocationId auto-resolves in single-location workspaces; no physical stock movement occurs yet; idempotency under idempotencyKey; and specific refusal reasons (invalid_transition, nothing_open, period_locked). This level of detail about side effects and constraints is exceptionally transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place. It opens with 'AGENT-FIRST' to signal priority, immediately states the core action, then packs critical behavioral constraints, refusal conditions, and idempotency into a compact, front-loaded structure. No superfluous wording; each clause adds necessary information. It is concise despite its length because it is efficiently organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, 0% schema coverage, and 7 parameters, the description provides a comprehensive picture: purpose, key parameter semantics, behavior (no physical movements), refusal reasons, and idempotency. It explains the primary side effect (creates a draft) and constraints. While it does not describe the return value, that is not critical for calling the tool. It covers all essential aspects an agent needs to invoke it correctly, including edge cases like locked periods and single-location resolution. It is complete for the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains receivedAt (ISO date, critical for period lock), defaultLocationId (optional, auto-resolve), and idempotencyKey (implicitly via idempotency). It does not explicitly describe workspaceId, poId, note, or expectedAt. However, workspaceId and poId are self-evident from naming and context, and note is trivial. The most non-obvious parameters are covered, leaving only expectedAt unexplained. This is a strong compensation given the schema's lack of descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Open a DRAFT Wareneingang against an open purchase order.' It specifies the action (open/create), the resource (draft goods receipt), and the context (against an open purchase order). It also differentiates from siblings by emphasizing 'Nothing physical happens yet' and that it is a draft creation, distinguishing it from posting, line editing, and reversal tools. The description is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides usage context: it creates a draft against an open purchase order, and it lists refusal conditions (when order not sent/received, no open quantity, locked period). However, it does not explicitly name alternative tools or state when to use this versus them (e.g., 'use goods_receipt_post to finalize'). It implies it is the initial step in the goods receipt workflow but lacks explicit exclusionary guidance. This gives clear context but leaves some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goods_receipt_getARead-only
One goods receipt with its header (including hasOverReceipt), its lines (quantity, unit-cost snapshot, location, lot/serial, inspection status, the over-delivered quantity, the stock movement each recognised line minted, and the reversal movement when it has been reversed) and its append-only decision trail. A foreign or unknown id is not_found, never cross-tenant data.
| Name | Required | Description | Default |
|---|---|---|---|
| grId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, and the description adds meaningful behavioral context: exact response contents, not_found for foreign/unknown ids, no cross-tenant leakage, reversal movement handling, and an append-only decision trail. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the primary resource, but it is one long parenthetical-heavy sentence that requires careful parsing. Every clause adds information, though better structuring would improve readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-resource getter with no output schema, the description is remarkably complete: it covers header fields, line-level fields, reversal behavior, decision trail, error semantics, and tenant isolation. Nothing essential for calling this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only indirectly clarifies params via 'foreign or unknown id' and 'never cross-tenant data'; it never explicitly names grId or workspaceId, but the implied mapping is reasonably inferable for a simple get-by-id tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb-like resource ('One goods receipt') and enumerates what is returned: header, lines, decision trail. The phrase 'One goods receipt' and the id/not_found semantics distinguish it from sibling list/preview/cancel tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to choose this tool versus goods_receipt_list, goods_receipt_preview, or goods_receipt_get_config. The description implies a single-receipt fetch by id, but it never states an exclusion or names alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goods_receipt_get_configARead-only
Read the workspace over-receipt posture: allowOverReceipt (default true, meaning an over-delivery is accepted and flagged rather than refused) and overReceiptPct (null means no cap; a number is an integer percentage of the ordered quantity above which a receipt is refused instead of flagged).
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces this with 'Read.' It adds meaningful behavioral context beyond the annotation by explaining the two returned values: allowOverReceipt's default and effect, and overReceiptPct's null/number semantics. This is valuable because there is no output schema to document these fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence with no filler and the main verb front-loaded. The parenthetical explanations are necessary to explain the config semantics, though the sentence is slightly heavy with nested clauses.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-config tool with one parameter and no output schema, the description covers the key return semantics and defaults, which is the main thing an agent needs to interpret the result. It is not exhaustive about error cases or response shape, but it is sufficient for correct invocation and basic interpretation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explicitly define workspaceId. However, 'workspace' in the description maps clearly to the workspaceId parameter, and the single required string parameter is self-explanatory in this context. The description does not add details about format or constraints, so it only partially compensates for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb and resource: "Read the workspace over-receipt posture." It also distinguishes itself from the sibling goods_receipt_set_config by framing this as a read operation, so an agent can tell the getter apart from the setter without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context that this is the read-only counterpart to configuration, and it is naturally used when inspecting over-receipt settings. It does not explicitly name alternatives or exclusions, but the verb 'Read' and the mention of workspace posture make the intended use obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goods_receipt_lines_for_matchARead-only
The receipt lines a three-way match may still bill for the given purchase-order lines: only recognised lines on a posted, non-reversed receipt whose billed quantity is below the received quantity. Each row carries the stable receipt-line id (the target I03 allocates landed cost to and I04 marks billed), the receipt number and date, the unit-cost snapshot and the open (unbilled) quantity.
| Name | Required | Description | Default |
|---|---|---|---|
| poLineIds | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true, so the agent knows this is a safe read. The description adds valuable behavioral context: it filters to posted, non-reversed receipts, only recognised lines, and open (unbilled) quantity. It also discloses that each row carries the stable receipt-line id, receipt number/date, unit-cost snapshot, and open quantity. This goes beyond the annotation and helps the agent understand what the result represents.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence that front-loads the core purpose and then adds precise filtering criteria and row contents. No wasted words; every clause adds meaning. It is long but information-dense, which is appropriate for a complex domain.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (three-way match domain, multiple filtering conditions), the description is quite complete. It explains the row contents and the selection criteria. It doesn't mention pagination, sorting, or whether the result is grouped by PO line, but with no output schema and a read-only annotation, the description covers the essential semantics. A small gap: it doesn't state whether the result is per PO line or per receipt line, though 'each row' implies per receipt line.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the meaning of the result rows but does not explicitly explain the two parameters (workspaceId, poLineIds). However, the description's phrase 'for the given purchase-order lines' clearly implies poLineIds is the input, and workspaceId is a standard scoping parameter. This is adequate but not fully explicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('list') and resource ('receipt lines for three-way match'), and precisely defines the scope: only recognised lines on a posted, non-reversed receipt whose billed quantity is below the received quantity. It also names the target I03/I04, which distinguishes it from sibling tools like match_three_way_list or goods_receipt_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool: when you need the receipt lines that a three-way match may still bill for given purchase-order lines. It doesn't explicitly state when not to use it or name alternatives, but the precise scoping and the mention of I03/I04 targets provide strong context. A small gap: no explicit exclusion like 'use match_three_way_list for existing matches'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goods_receipt_listARead-only
List goods receipts with filters: status (one value or an array of draft | posted | reversed | cancelled), poId, supplierContactId, a receivedAt date range (fromDate / toDate), hasOverReceipt (the exception cut: only the deliveries that came in over what was ordered) and a free-text search q over the number and the note. Each row carries the order number, the supplier name, the line count, the receipt value in Rappen and the over-delivered quantity. Accepts a savedViewId (G00 saved-view seam). Newest received date first.
| Name | Required | Description | Default |
|---|---|---|---|
| q | No | ||
| poId | No | ||
| status | No | ||
| toDate | No | ||
| fromDate | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| hasOverReceipt | No | ||
| supplierContactId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool with readOnlyHint=true, so the description's main contribution is the output row structure (order number, supplier, line count, value in Rappen, over-delivered quantity) and the sort order (newest first). It adds useful behavior info without contradicting the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences cover the tool's purpose, all key filters, output row content, and sort order without redundancy. The most important information is front-loaded, and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with nine parameters and no output schema, the description provides solid coverage of filters, output fields, and sorting. The only notable omission is the required workspaceId, though it is likely standard across workspace-scoped tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: it explains status values, poId, supplierContactId, date range, hasOverReceipt, q, and savedViewId. It omits workspaceId (the only required parameter), but the rest are well covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'List goods receipts' – a specific verb and resource – then enumerates the actual filters, making it unmistakably clear what the tool does. It differentiates from siblings like goods_receipt_cancel and goods_receipt_create by being the listing variant.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for retrieving goods receipts, but it never explicitly states when to choose this over other list tools or when not to use it. No alternative tools or exclusion criteria are mentioned, so the agent is left to infer the appropriate context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goods_receipt_postA
Post the goods receipt: in ONE transaction re-validate every line against the LIVE purchase order, write the J02 stock movements (movementType receipt, the order line CHF base cost as the snapshot, sourceDocumentType goods_receipt), append the shared received-quantity trail, raise po_line.received_qty, and freeze the document. There is deliberately NO date parameter: the movements carry the receipt own receivedAt and the period lock is asserted against that same date. An over-delivery is ACCEPTED at its full quantity by default and RECORDED as an exception (the excess lands on the line as overReceiptQty, the header carries hasOverReceipt, and an append-only over_receipt event names who took how many extra units); it is never silently clamped. A workspace that wants a hard ceiling gets one: with allowOverReceipt false anything above the open quantity is refused with qty_exceeds_open, and with an overReceiptPct set anything above the tolerance is refused with over_receipt, and above the ceiling the WHOLE receipt rolls back. A line whose item is not stock-tracked advances the ordered quantity and mints no movement. Any guard trip rolls the WHOLE receipt back: no movement, no trail row, no quantity change. A replay under the same idempotencyKey returns the posted document and writes nothing; a second post under a different key is invalid_transition.
| Name | Required | Description | Default |
|---|---|---|---|
| grId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully explains the write effects (J02 movements, trail, received_qty, freeze), exceptions (over-delivery handling), rollback atomicity, and idempotency. This goes far beyond what any annotation could provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but every sentence adds a behavior or constraint. Front-loaded with the core purpose and then details, no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (atomicity, exceptions, config, idempotency), the description covers the essential behavior, error types, and return semantics for replay. It lacks an explicit success return description, but the replay mention implies the posted document is returned.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no parameter descriptions (0% coverage), but the description explains idempotencyKey (replay behavior), workspace (per-workspace config), and grId (the goods receipt being posted). It does not use the exact parameter names, but the mapping is inferable. Optional properties like allowOverReceipt are mentioned as config but not explicitly tied to call parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (post) and resource (goods receipt) with a clear action. Distinguishes from siblings like goods_receipt_cancel and goods_receipt_create by focusing on the posting step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides rich context on the post operation's behavior, including config-dependent over-delivery handling and idempotency semantics. However, it doesn't explicitly state preconditions like when a receipt is ready to post or how it differs from goods_receipt_accept_lines. It implies it is the final commit step but doesn't name alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goods_receipt_previewARead-only
PURE impact preview of a draft goods receipt: per line the ordered, already-received, open and ceiling quantities (ceiling null means over-delivery is uncapped), the proposed and resulting received quantity, the unit-cost snapshot, whether the line moves stock, the over-delivered quantity it would record (overReceipt / overReceiptQty), and the exact rejection codes a post would refuse with: per line the quantity, reference, location and J01 lot/serial refusals (qty_exceeds_open, over_receipt, invalid_qty for a serial line of more than one unit, invalid_reference, location_required, lot_required, serial_required, lot_item_mismatch, serial_item_mismatch), and on the top-level document issues array the period_locked one, which is not a property of any single line. An over-delivery WITHIN the policy is reported but is NOT an issue, because it posts. Plus the total value in Rappen and a postable flag that is false when either array is non-empty. Writes nothing and is safe to call repeatedly.
| Name | Required | Description | Default |
|---|---|---|---|
| grId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the readOnlyHint annotation by detailing exact outputs, including rejection codes, edge cases (over-delivery within policy is not an issue), and explicitly states 'Writes nothing and is safe to call repeatedly.' This provides comprehensive behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense paragraph with many details, making it hard to parse. It is thorough but not concise or well-structured, lacking bullet points or clear separation of concerns.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description exhaustively covers the output and edge cases, which is critical given the lack of an output schema. However, it omits parameter explanations, which reduces completeness for a tool with two required inputs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, and the description never mentions workspaceId or grId, leaving their purpose unexplained. For a tool with only two parameters, this is a significant gap that the description does not compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states it is a 'PURE impact preview of a draft goods receipt' and enumerates the exact data returned, including quantities, rejection codes, and postable flag. This clearly identifies the tool's function and distinguishes it from mutation siblings like goods_receipt_post or goods_receipt_reject_lines.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for previewing before posting by mentioning the postable flag and rejection codes, but it does not explicitly state when to use this tool instead of alternatives. No when/when-not guidance is provided, though the read-only nature is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goods_receipt_reject_linesA
Refuse HELD lines after a quality check. No stock movement and no order quantity: a rejected quantity never entered the warehouse, so the ordered quantity simply stays open and can be received again on a later delivery. reason is required and is recorded with the actor on an append-only rejection event. A line that is not pending is invalid_transition.
| Name | Required | Description | Default |
|---|---|---|---|
| grId | Yes | ||
| reason | Yes | ||
| lineIds | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so thoroughly. It states that rejection causes no stock movement and no order quantity change, that the ordered quantity remains open and can be received later, that the reason is recorded with the actor on an append-only event, and that a non-pending line triggers an invalid_transition. These are precisely the side effects and failure conditions an agent needs to know, and they are all present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each carrying essential information. It opens with the core purpose, then explains the consequences, and finally states the constraint on line state. There is no redundancy or fluff. The structure is logical and front-loaded, making it easy for an agent to quickly grasp the key facts.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating tool with no annotations and no output schema, the description covers the essential context: when to use it, what effects it has on stock and order quantity, the requirement for a reason, the append-only event, and the invalid_transition condition. It does not describe the return value or how to obtain grId/lineIds, but these are likely external concerns. Overall, the description is sufficiently complete for an agent to call the tool correctly, with minor gaps around parameter acquisition and response handling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no parameter descriptions (0% coverage), so the description must compensate. It explains that 'reason' is required and recorded with the actor on an append-only event, and it implies that 'lineIds' refers to the lines being rejected and that they must be pending. However, it does not clarify the meaning of 'grId', 'workspaceId', or 'idempotencyKey'. While some are self-evident from context, the description does not fully cover all parameters, leaving some ambiguity for an agent unfamiliar with the domain.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: refuse HELD lines after a quality check. It specifies the action (refuse/reject), the resource (lines on a goods receipt), and the condition (HELD). It also differentiates from related tools like goods_receipt_accept_lines by explicitly contrasting with acceptance behavior (no stock movement, quantity stays open). The verb 'refuse' and the domain context make the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context: it is to be used 'after a quality check' on HELD lines, and it explains that a non-pending line results in an invalid_transition, which guides when the tool is applicable. However, it does not explicitly name alternative tools (e.g., goods_receipt_accept_lines) or state when to use those instead. The guidance is implied rather than explicit, but still adequate for an agent to choose correctly in most scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goods_receipt_reverseA
Storniere einen gebuchten Wareneingang: the ONE correction for a posted receipt. Nothing about the original document, its movements or its trail rows is edited or deleted. What is written is the compensation: a J02 return movement of equal magnitude and opposite sign per recognised line (same date, same source-document link, so the pair is findable together), a negative trail row so the received-quantity trail still sums to po_line.received_qty, the received_qty rollback, and the original marked reversed with who and when. Refused with line_already_billed when a three-way match has already billed a line, with period_locked when the receipt own period is sealed, and with insufficient_stock when the goods have since left the warehouse (in which case the honest correction is an inventory adjustment, not an un-receipt). reason is required. Idempotent under idempotencyKey.
| Name | Required | Description | Default |
|---|---|---|---|
| grId | Yes | ||
| reason | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations to rely on, the description fully carries the behavioral burden. It explicitly states nothing is edited or deleted ('Nothing about the original document, its movements or its trail rows is edited or deleted'), describes the compensating J02 movement, negative trail row, rollback, and marking, and lists error conditions. This is exceptionally transparent for a mutating operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence adds value: purpose first, then non-destructive behavior, movement details, trail accounting, refusal conditions, and required fields. It is structured to front-load the core action and then progressively details side effects and constraints without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lacking output schema and annotations, the description covers all necessary operational context: what the tool does, what it does not do, the exact accounting movements created, trail integrity, error scenarios, the required reason field, and idempotency guarantees. An agent has everything needed to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains 'reason is required' and clarifies idempotencyKey behavior ('Idempotent under idempotencyKey'). It does not explicitly describe grId or workspaceId, but those are strongly inferred by the tool name and standard conventionscribe. It adds meaningful semantics beyond the bare schema for the most important parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear action: 'Storniere einen gebuchten Wareneingang' (reverse a posted goods receipt), and then explicitly calls it 'the ONE correction for a posted receipt.' This distinguishes it from related tools like goods_receipt_cancel and clarifies the resource and verb precisely.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides strong usage guidance by stating this is the single correct tool for reversing a posted receipt, and gives concrete refusals (line_already_billed, period_locked, insufficient_stock) with an explicit alternative for the insufficient_stock case: 'the honest correction is an inventory adjustment, not an un-receipt.' This tells the agent when not to use it and what to use instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goods_receipt_set_configA
Set the workspace over-receipt posture. The DEFAULT (allowOverReceipt true, no overReceiptPct) accepts an over-delivery at its full quantity and flags it on the receipt, because blocking it at the loading dock does not un-deliver the goods, it only stops the ledger from saying they arrived. allowOverReceipt false is the hard ceiling: anything above the open order quantity is refused with qty_exceeds_open. OMIT overReceiptPct for no cap; set it (1..100, an integer percentage of the ORDERED quantity, rounded DOWN so a tolerance is never wider than granted) to refuse above that tolerance with over_receipt instead of flagging. Plain policy: posts no journal entry, moves no stock and never rewrites history, so changing it affects only what a future post will accept.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| overReceiptPct | No | ||
| allowOverReceipt | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers richly: it discloses the operation is side-effect-free ('posts no journal entry, moves no stock and never rewrites history'), explains the rationale for the default posture, specifies exact refusal modes and the exact error codes (qty_exceeds_open, over_receipt), and states the rounding direction so the agent understands exact boundary behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: purpose, default semantics, refusal behavior, tolerance cap, and side-effect scope are all dense, non-redundant information. Slightly long as one unbroken paragraph — bullet structure would improve scannability — but there is zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an unannotated, schema-poor, no-output-schema config tool, the description covers nearly everything an agent needs: all config modes, error codes, side-effect scope, and rationale. Small gaps remain — workspaceId and idempotencyKey semantics are left to inference, and the return value is never described — minor for a setter tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It thoroughly explains allowOverReceipt and overReceiptPct including the 1..100 integer range and rounding-down guarantee. The two remaining params (workspaceId, idempotencyKey) are not mentioned, though their roles are largely self-evident from their names. Near-complete compensation for a sparse schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource: 'Set the workspace over-receipt posture.' It then differentiates itself from the goods_receipt_* siblings by detailing exactly which config dimension it controls (over-receipt acceptance vs refusal) and its error codes, leaving no ambiguity against goods_receipt_create/post/get or inventory_set_config.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Strong internal guidance on which configuration mode to use: default accept-and-flag, hard ceiling when allowOverReceipt=false, and tolerance cap via overReceiptPct. However, it never explicitly names alternatives or says 'use X instead' — an agent must infer that goods_receipt_get_config is the read counterpart. Clear context but no explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goods_receipt_upsert_linesA
Add, change or remove lines on a DRAFT goods receipt. Each op is { op: add | change | remove, poLineId (add), lineId (change/remove), qty (> 0, whole units), locationId?, lotId?, serialId?, unitCostRappen? (defaults to the order line CHF base price), inspectionStatus? (none | pending), note? }. A line marked pending is HELD at post: no stock movement and no order quantity until goods_receipt_accept_lines releases it. accepted and rejected are decisions with an actor attached and are reachable only through the accept / reject verbs, never as a data field. Editing a posted receipt is refused (invalid_transition): the correction for a posted mistake is goods_receipt_reverse, never an edit.
| Name | Required | Description | Default |
|---|---|---|---|
| ops | Yes | ||
| grId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it does well: it explains the pending state ('HELD at post: no stock movement and no order quantity until goods_receipt_accept_lines releases it'), the refusal on posted receipts, and the distinction between decisions and data. It doesn't cover every edge case (e.g., success response or idempotency behavior), but the core operational semantics are clearly conveyed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but not bloated. It front-loads the purpose and then packs the op specification in a compact structured notation. Each sentence serves a distinct explanatory purpose (ops format, pending behavior, transition restrictions). It could be trimmed slightly without loss, but it remains efficient and well organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the sparse schema (no item schema for ops, no output schema, no enums), the description is remarkably complete for a complex tool. It explains the ops semantics, constraints, and key behavioral nuances. It leaves minor gaps like what a successful call returns or whether the ops array must be non-empty, but these are not critical for invocation and the description already exceeds typical expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the ops parameter is only typed as an array with no item schema, so the description must compensate. It does thoroughly: it specifies the operation structure (op type, required fields per op, constraints on qty, optional fields with defaults like unitCostRappen and inspectionStatus). The other parameters (workspaceId, grId, idempotencyKey) are standard and don't require elaboration. This description adds substantial meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Add, change or remove lines') and a specific resource ('a DRAFT goods receipt'), and immediately distinguishes it from related operations by naming the constraint that it only applies to drafts. It also references sibling tools like goods_receipt_accept_lines and goods_receipt_reverse, making the tool's scope unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to use the tool (on a DRAFT receipt) and when not to use it ('Editing a posted receipt is refused'), and directs the user to the correct alternative (goods_receipt_reverse). It also clarifies that accepted/rejected decisions are not data fields but are handled via accept/reject verbs, preventing misuse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
go_productiveA
Promote a verified Testmandant to real books, in place, with NO re-import: the workspace simply stops being a Testmandant (kind = live), every id and timestamp stable, the A03 audit chain continuous (OR Art. 957a). The least-reversible act in the product and irreversible: it re-runs the WHOLE Eröffnungsprüfung one final time and refuses unless every control passes and none diverges (check_not_clean), no row traces to a demo seed (sandbox_contains_demo_rows), the company profile is real (needs_company_profile), and no live workspace already holds the UID (live_workspace_exists). Requires promote_workspace AND commit_migration. confirmedName is a type-to-confirm against the company legal name, checked engine-side so the agent face cannot skip it; an agent caller is draft-staged until a human confirms. Promoting an already-live workspace is a no-op, never a double-promote. CONSEQUENCE: Promotes the trial workspace to real books in place; the cutover is irreversible.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| workspaceId | Yes | ||
| confirmedName | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full disclosure burden and does so thoroughly: it exposes irreversibility, the final full audit re-run, refusal conditions, draft-staging for agent callers, and no double-promotion behavior. This is far more transparent than typical write operations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the core action up front, and most sentences carry real information. However, the final CONSEQUENCE sentence largely repeats earlier statements such as 'in place' and 'irreversible', so not every sentence fully earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an irreversible, high-stakes operation with no annotations and no output schema, the description covers prerequisites, refusal checks, permissions, human-confirm staging, and no-op behavior. Missing explicit meaning for three of the four parameters and any result-shape guidance, but the overall decision context is unusually complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but only confirmedName is explained (type-to-confirm against legal name). workspaceId, planId, and idempotencyKey semantics are left entirely to inference, with no explanation of how idempotency or planId behave.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the exact action with a specific verb and resource: 'Promote a verified Testmandant to real books, in place, with NO re-import' and clarifies the core state change (kind = live, stable ids, continuous audit chain). This distinguishes it clearly from migration and re-import siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states explicit preconditions: required permissions, the checks that must pass, the confirmedName type-to-confirm requirement, and that promoting an already-live workspace is a no-op. It does not name an alternative sibling tool, but it gives enough context and exclusion conditions for an agent to know when this is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hr_absence_cancelA
Cancel an absence by flipping its status to cancelled (append-only spirit: the row stays, it is never deleted). Requires hr.manage.
| Name | Required | Description | Default |
|---|---|---|---|
| absenceId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly states that the row is never deleted, only its status changes, which is a key behavioral trait. It also mentions the required permission. However, it does not disclose side effects on related data, idempotency behavior, or error handling, but the core mutation semantics are clearly conveyed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with a parenthetical that adds essential context without padding. The core action is front-loaded, and the permission requirement is placed at the end, maintaining clarity and efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has three required parameters, no output schema, and no annotations. The description covers the core operation and permission but omits details like return value, idempotency semantics (despite the idempotencyKey parameter), and handling of already-cancelled absences. For a simple cancellation tool, this is acceptable but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining the three parameters (workspaceId, absenceId, idempotencyKey). It does not mention any of them. The parameter names are self-explanatory to a human, but the description adds no meaning beyond the bare schema, leaving an agent without guidance on how to construct valid calls.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Cancel an absence by flipping its status to cancelled'. It names the resource (absence) and the specific operation (cancel), and distinguishes it from siblings like hr_absence_list or hr_absence_record. The append-only clarification further differentiates it from delete operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context (cancelling an absence) and mentions the required permission (hr.manage). It does not explicitly name alternatives or exclusions, but the purpose is unambiguous enough that an agent can infer when to use this tool versus related ones like hr_absence_record or hr_absence_list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hr_absence_listARead-only
List absences, SELF-SCOPED: without hr.manage the caller receives only their own linked employees rows, because sick leave is health data (revDSG Art. 5 lit. c Ziff. 2). employeeId, from and to narrow within scope; they can never widen it to a colleagues records.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| employeeId | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide readOnlyHint=true, so the description adds meaningful behavioral context: role-based self-scoping, the legal privacy basis, and the guarantee that parameters cannot widen access to other employees' records.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two focused sentences front-load the action and the most important behavioral constraint. The legal citation is relevant and not padding; every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, permission boundary, and filter behavior well, but the five-parameter schema includes workspaceId (required) and savedViewId that receive no explanation, and there is no output schema or return-shape description. Adequate for simple listing but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate. It explains employeeId, from, and to as scope-narrowing filters, but it does not explain the required workspaceId or savedViewId, nor does it specify expected formats or value constraints. This is only partial compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('List absences') and resource, and the 'SELF-SCOPED' qualifier immediately distinguishes it from other absence-related tools. The health-data and employee-rows explanation reinforces what kind of data this returns and who can see it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: callers without hr.manage see only their own linked employee's absences, and employeeId/from/to can only narrow results. It does not explicitly name alternative sibling tools, but the scope rule is enough to guide correct use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hr_absence_recordB
Record an absence (ArG Art. 46 record-keeping): kind is vacation|sick|other, plus from_date and to_date (YYYY-MM-DD). Requires hr.manage. An overlapping absence for the same employee is ACCEPTED with overlap:true in the result (a half-day sick during vacation), never merged. Sick leave is health data (revDSG Art. 5 lit. c Ziff. 2).
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | ||
| notes | No | ||
| toDate | Yes | ||
| fromDate | Yes | ||
| employeeId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full behavioral burden and does well: it discloses that overlapping absences are accepted rather than merged, that the result includes overlap:true, and that sick leave constitutes sensitive health data. It could additionally explain idempotency behavior and what a successful creation returns, but the disclosed edge case is significant and valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the main purpose, then supplies permission, behavioral edge case, and data sensitivity context in three sentences. The legal citations add mild extra load but are relevant and do not obscure the core meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description covers the most important behavior, permission, and overlap handling, making the tool usable. It is incomplete on idempotency semantics and the full expected response shape, which an agent would need to invoke it confidently across retry or partial-success scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the 7 parameters. It does explain kind's allowed values and the date format, but it refers to from_date and to_date using snake_case while the schema declares fromDate and toDate, creating a potential mapping error. It also omits semantics for idempotencyKey, notes, workspaceId, and employeeId.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Record an absence'), the resource, and the core fields with allowed kind values and date format. It differentiates itself from the sibling tools hr_absence_list and hr_absence_cancel by using an explicit write action and naming obligations, though it does not directly name those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a prerequisite ('Requires hr.manage') and the legal record-keeping context, but it does not explain when to use this tool versus hr_absence_list or hr_absence_cancel, nor does it provide when-not or alternative conditions. The reader has to infer usage boundaries from the tool name and sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hr_employee_getARead-only
Read one employee. The AHV number is returned only when includeSensitive is true AND the caller holds hr.sensitive; otherwise it is masked (756-...) and ahvRestricted is true.
| Name | Required | Description | Default |
|---|---|---|---|
| employeeId | Yes | ||
| workspaceId | Yes | ||
| includeSensitive | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations include readOnlyHint=true, and the description adds valuable conditional behavior details: AHV masking logic with includeSensitive flag and hr.sensitive permission, plus the ahvRestricted field. This goes beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, highly efficient: first states core purpose, second conveys the sensitive-data behavior. No fluff, information dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-one resource with no output schema and clear params, the description is nearly complete. It covers the main behavioral nuance (AHV masking) and the permission required. Missing return format/error conditions are minor for a simple read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains the effect of includeSensitive on the response (AHV masking) and the permission requirement. It does not elaborate on workspaceId or employeeId, but their meaning is clear from names and required status.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it reads one employee, distinguishing from hr_employee_list (list) and hr_employee_upsert (write). The conditional masking behavior is specific and descriptive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for reading a single employee but provides no explicit when-to-use guidance or comparison with hr_employee_list. It does not state alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hr_employee_listARead-only
List employees (the roster). AHV numbers are masked unless includeSensitive is true and the caller holds hr.sensitive. includeArchived includes ended/archived people.
| Name | Required | Description | Default |
|---|---|---|---|
| savedViewId | No | ||
| workspaceId | Yes | ||
| includeArchived | No | ||
| includeSensitive | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description adds meaningful behavioral detail: AHV numbers are masked unless includeSensitive is true AND the caller holds hr.sensitive; includeArchived includes ended/archived people. This tells the agent about permission requirements and data visibility that annotations alone don't convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences; first is a crisp definition, second packs the two key behavioral modifiers. No redundant words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list operation with readOnlyHint and no output schema, the description covers the core behavior and the sensitive/archived edge cases. It is somewhat incomplete regarding savedViewId semantics and does not mention pagination or response shape, but these are minor for a simple roster listing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 0% description coverage, so the description carries the burden for parameter meaning. It explains includeSensitive (requires hr.sensitive for unmasked AHV) and includeArchived (includes ended/archived people), which is valuable. However, the required workspaceId and especially savedViewId are left unexplained, so an agent may not know what savedViewId controls.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'List employees (the roster)'. The parenthetical 'the roster' clarifies that this returns the full set of employees rather than a single employee, distinguishing it from sibling hr_employee_get and hr_employee_upsert.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies using this tool to read the employee roster, but it does not explicitly contrast with alternatives such as hr_employee_get for a single employee or hr_employee_upsert for changes, nor does it say when not to use it. It gives context via includeArchived/includeSensitive but no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hr_employee_upsertA
Add or edit an employee (Mitarbeitende): first/last name, employment_pct (integer 1-100), starts_on and optional ends_on, an optional linked contact (the counterparty a reimbursement pays), and an optional actor_ref (the member whose calls self-scope to this person). Requires hr.manage. ahv_nr is OPTIONAL and captured only for the external payroll hand-off (revDSG Art. 6 minimisation); writing it additionally requires hr.sensitive, and it is masked on every read that does not hold hr.sensitive. Pass employee.id to edit an existing row.
| Name | Required | Description | Default |
|---|---|---|---|
| employee | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses permission requirements, data minimization (revDSG Art. 6), masking of ahv_nr on reads without hr.sensitive, and optional field behavior. However, it does not explain the required idempotencyKey's behavior or what the response looks like, so it is not fully exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but each sentence earns its place, covering action, fields, permissions, privacy, and edit semantics. It is front-loaded with the core purpose and does not waste words, though a more structured list might improve scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a nested-object mutation with no annotations and no output schema, the description covers the main fields and critical constraints well. But it leaves out idempotency semantics, workspaceId purpose, and any expectation of return values or failure modes, so an agent would still have significant unknowns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema is bare (0% coverage), and the description adds substantial meaning for the nested employee object: it explains employment_pct, starts_on, ends_on, contact, actor_ref, ahv_nr, and employee.id. However, required top-level parameters workspaceId and idempotencyKey are never explained, leaving gaps for the agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Add or edit an employee' and lists concrete fields, clearly establishing a write/upsert operation. It also differentiates the edit case with 'Pass employee.id to edit an existing row,' which distinguishes this tool from read-only siblings like hr_employee_list and hr_employee_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to create vs edit (employee.id present) and mentions required permissions (hr.manage, hr.sensitive for ahv_nr). It does not explicitly name alternative tools or exclusion conditions, but the upsert nature and edit instruction make usage reasonably clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
implementation_decision_recordA
Record one decision in the append-only implementation decision log (OR 957a Nachprüfbarkeit applied to the implementation itself): title, context and the decision. There is no update path. Gates on manage_implementation.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| context | No | ||
| decision | Yes | ||
| projectId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It reveals the key write semantics: append-only, no update path, and permission gating. It does not describe idempotency behavior or return values, but the most important behavioral constraints for using a write tool are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, purpose-first, with no filler. Each sentence contributes distinct information: the action and resource, the append-only/no-update constraint, and the permission gate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple append-only write tool, the description covers purpose, content fields, update behavior, and access control. It does not describe the response or idempotency duplicate handling, but those are not essential given the straightforward schema and no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for undocumented parameters. It explains the content fields 'title, context and the decision,' which adds meaning beyond the bare schema. However, it does not clarify workspaceId, projectId, or idempotencyKey, although those parameter names are fairly self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Record one decision') and a specific resource ('append-only implementation decision log'), and adds the critical append-only trait. This distinguishes it from siblings like implementation_signoff_record and implementation_task_set without requiring schema inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly frames when to use the tool: for recording new decisions in an append-only log, with the explicit exclusion 'There is no update path.' It also mentions the permission gate ('Gates on manage_implementation'). It does not explicitly name sibling alternatives, but the append-only decision-log framing provides adequate context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
implementation_parallel_checkA
Compute TILL’s own figures for a period from the OWNING read models (A08 trial balance, A07 return per Ziffer, A16/A17 open items) and compare, per figure, to the declared figures: { kind, ref, declaredRappen, computedRappen, differenceRappen, status }. ZERO tolerance, no dial: a 0-Rappen difference passes, ANY nonzero fails. A figure with no declaration is not_asserted, never green. Re-running with changed evidence voids any bound tieout / parallel_run_close sign-off. Gates on manage_implementation.
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| projectId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since no annotations are provided, the description carries the full burden. It discloses key behavioral traits: the tolerance rule (zero tolerance), the handling of undeclared figures (not_asserted, never green), and the side effect of re-running (voiding tieout/parallel_run_close sign-off). It does not mention authorization, rate limits, or reversibility of the operation, but for a read/comparison tool these are less critical. It is fairly transparent given no annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-rich, with each sentence serving a purpose: it states the operation, the sources, the output format, the tolerance policy, the handling of missing declarations, and the side effect of re-running. It is front-loaded with the core purpose and avoids fluff. It is structured logically and not overly verbose for the complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool involving computation and comparison, the description covers the main aspects: what it does, the output structure, tolerance rules, missing-declaration behavior, and re-run side effects. It does not explicitly mention prerequisites (e.g., that declarations must exist or that read models are populated), but it is fairly complete. Since there is no output schema, the inclusion of the output format is helpful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. However, it does not explain workspaceId, projectId, period, or idempotencyKey beyond mentioning 'period' in the context of the computation. It does not define period format, idempotencyKey usage, or how workspace/project are used. This is a significant gap since the schema provides no descriptions either.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: it computes TILL's own figures from specific read models (A08, A07, A16/A17) and compares them to declared figures. It names the exact sources and the comparison action, and it provides the output shape. This distinguishes it from siblings like implementation_parallel_declare (which declares figures) and implementation_parallel_status (which likely just returns status).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit conditions: ZERO tolerance, a 0-Rappen difference passes, ANY nonzero fails, and a figure with no declaration is not_asserted. It also warns that re-running with changed evidence voids sign-offs. However, it does not explicitly name alternative tools or state 'use this instead of X', though the context implies this is the check operation vs. declaration or status. It does mention a gate on manage_implementation, which is a precondition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
implementation_parallel_declareA
Declare the prior system’s figures for one filing period (entered exactly like G11 control totals: declared, hashed, never edited in place; a correction is a new declaration superseding the old, both retained). figures is an array of { kind (trial_balance|vat_return|open_items_ar|open_items_ap), ref, declaredRappen (integer) }. The period must align to the method boundary (effektiv quarter YYYY-Qn, saldo semester YYYY-Hn) or refuses (period_outside_window). A cutover combined with a method change refuses without a recorded mwst_method sign-off (method_change_needs_signoff). Gates on manage_implementation.
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| figures | Yes | ||
| projectId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does so thoroughly: it discloses immutability, hashing, correction semantics, retention of superseded declarations, validation refusals, the sign-off dependency, and access gating. This is more behavioral context than most tool descriptions provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The information is dense and relevant, but it is compressed into a long, semicolon-laden paragraph without clear structure. The purpose is front-loaded, yet terms like 'G11 control totals', 'effektiv', and 'saldo' make it harder to parse than necessary; better formatting would improve scannability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is complex, with five required parameters, no output schema, and no annotations. The description covers the non-obvious business rules, validations, and authorization gate. It omits return-value behavior and explicit idempotency-key semantics, but the core call context is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds real meaning for 'figures' (kinds, ref, declaredRappen) and 'period' (YYYY-Qn/YYYY-Hn formats), but leaves workspaceId, projectId, and idempotencyKey unexplained. Given the 0% schema coverage, the gaps keep this at a mid score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific verb ('declare') and resource ('prior system's figures for one filing period'), and distinguishes the operation from related implementation and migration tools by emphasizing the parallel-implementation context and the G11 control-total analogy. The scope is precise enough that an agent can tell it apart from siblings like implementation_parallel_status or migration_declare_control_total.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It establishes clear context and prerequisites: declarations are immutable, corrections require a new declaration, the period must match the method boundary, a method-change cutover requires mwst_method sign-off, and the operation is gated on manage_implementation. It does not explicitly name alternative tools or state when not to use it, so it stops short of a full when-vs-alternatives guide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
implementation_parallel_statusARead-only
Read the parallel-run status: per declared period the latest check’s roll-up (passed | failed | not_asserted) and an overall status that is never green while a declaration or a passing check is missing (G11 three-status honesty). Gates on manage_implementation.
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true, and the description's 'Read' is consistent with that (no contradiction). The description adds genuine behavioral value beyond the annotation: the overall status is 'never green while a declaration or a passing check is missing', the three-status honesty rule (G11), and the permission gate on manage_implementation. This meaningfully clarifies what the status actually means.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose and return shape, with the gating detail at the end. Every clause earns its place. Minor deduction for dense domain jargon ('G11 three-status honesty', 'declared period') that goes unexplained and could confuse an agent unfamiliar with the implementation domain.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only status tool with two standard params and no output schema, the description covers the return semantics (both the per-period roll-up and the overall status rule) and the auth gate. It does not define what a 'declared period' is or state the output shape explicitly, but a status-read tool of this simplicity is nearly fully specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does nothing to explain workspaceId/projectId or how they relate to 'per declared period'. However, both parameter names are highly conventional across this API and self-explanatory in context, so the compensation gap is low-risk. The names carry the meaning rather than the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb-resource pair ('Read the parallel-run status') and then enumerates exactly what is returned: per-period latest check roll-up (passed | failed | not_asserted) and an overall status. The read semantics and the two status layers clearly distinguish it from siblings like implementation_parallel_declare and implementation_parallel_check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides context that this tool 'gates on manage_implementation', which implies it is a permissioned status read within an implementation flow. However, it never explicitly contrasts itself with the closely related siblings implementation_parallel_check, implementation_parallel_declare, or implementation_project_get, nor states when to prefer this over them. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
implementation_project_closeA
Close the implementation project (confirm-gated). Legal only from the live (Stabilisierung) phase, with the Stabilisierung and parallel-run tasks terminal, and the parallel_run_close and source_cancellation sign-offs recorded (or their canon tasks not_applicable with a reason). Each open leg refuses in isolation, naming what is open. Denylisted from automation (a rule engine must not close). Gates on manage_implementation.
| Name | Required | Description | Default |
|---|---|---|---|
| confirmed | Yes | ||
| projectId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description bears full responsibility. It discloses confirm-gated behavior, refusal with reasons when legs are open, automation denylist, and a permission gate on manage_implementation. However, it does not state the success effect (e.g., project transitions to closed) or idempotency semantics despite the idempotencyKey parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a dense four‑sentence paragraph that front‑loads the core action. It packs a lot of precondition information into a compact space, though the specialized jargon (Stabilisierung, parallel_run_close) increases cognitive load without sacrificing brevity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity-decided preconditions, the description covers the main gates and restrictions well. However, it omits parameter semantics and does not describe what happens on success or what errors look like beyond the refusal behavior. Since no output schema is present, these gaps leave the agent with incomplete information for a mutating operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description fails to explain the parameters. It only indirectly hints at the 'confirmed' parameter via 'confirm-gated', while workspaceId, projectId, and idempotencyKey are never described. With no schema descriptions, this is a critical gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Close' and the resource 'implementation project', and adds 'confirm-gated' to specify the operation's nature. It distinguishes itself from the many sibling implementation_* tools by being the closure action with explicit conditions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly defines when the tool is legal: from the live (Stabilisierung) phase, with Stabilisierung and parallel-run tasks terminal, and sign-offs recorded or canon tasks not_applicable. It also states when it must not be used: denylisted from automation span> managed by a rule engine. These are firm when/when-not conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
implementation_project_createA
Create the implementation project for this workspace: the cutover as a first-class object above the migration plan(s). Takes the source system, the target cutover date (the fixed anchor every reconciliation figure is declared against), an optional freeze window, the MWST method (effektiv | saldo) and methodChange:true when the cutover is combined with a method change (which then needs an mwst_method sign-off before any parallel figure). Refuses a cutover date in the past (cutover_in_past), an inverted freeze window, and a second open project on the workspace (project_already_open). Gates on manage_implementation.
| Name | Required | Description | Default |
|---|---|---|---|
| freezeEnd | No | ||
| mwstMethod | Yes | ||
| cutoverDate | Yes | ||
| freezeStart | No | ||
| workspaceId | Yes | ||
| methodChange | No | ||
| sourceSystem | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well: it discloses refusal conditions (cutover_in_past, inverted freeze window, project_already_open) and the authorization gate. It also explains the methodChange implication requiring an mwst_method sign-off before any parallel figure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the purpose and every sentence adds value: parameter semantics, validation failures, and the gate. It is dense and slightly run-on, but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters, no annotations, and no output schema, this description covers a lot of ground: purpose, parameter semantics, failure modes, and permissions. Remaining gaps such as idempotencyKey behavior and return value meaning keep it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must add parameter meaning, and it does for most parameters: sourceSystem, cutoverDate as the fixed anchor, mwstMethod values, the freeze window, and methodChange semantics. It omits the required idempotencyKey parameter, which prevents a perfect score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Create the implementation project for this workspace', and adds domain context by defining the cutover as a first-class object above migration plan(s). This differentiates it from migration_create_plan and from implementation_project_get/list/close in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context is clear: it is clear this tool is for creating the workspace-level implementation project, and it states a key precondition and gate with 'Gates on manage_implementation' and the refusal on a 'second open project'. It does not explicitly name alternatives, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
implementation_project_getARead-only
Read the implementation project end to end: the evidence-derived phase (discovery -> extraction -> mapping -> rehearsal -> cutover -> parallel_run -> live -> closed; live is the Stabilisierung), the one next action, the blocked-first task list, the sign-off (Freigabe) rows, the decision log, the parallel-run status, and the load-bearing computed statuses tieout.passed / parallel.passed, each true ONLY when the deterministic checks are clean AND the bound human sign-off stands. Gates on manage_implementation.
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, so the agent knows it's safe. The description adds useful behavioral context: the computed statuses (tieout.passed / parallel.passed) are 'true ONLY when the deterministic checks are clean AND the bound human sign-off stands' – this is a non-obvious semantic that helps the agent interpret results. No contradiction, but it doesn't disclose return format or pagination.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense with information but well-structured. It front-loads the core purpose ('Read the implementation project end to end') and then enumerates components. It's a long sentence, but every element adds value. The final 'Gates on manage_implementation' is slightly cryptic but concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (many components and derived statuses) and no output schema, the description covers the major features thoroughly. It explains the phase lifecycle, the status logic, and highlights important nuances (live is Stabilisierung). It omits details about response format but for a read tool, that's acceptable. WorkspaceId/projectId are enough to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% – the schema has projectId, savedViewId, workspaceId but no descriptions. The tool description doesn't explain parameters either. However, the parameters are self-explanatory (workspaceId, projectId are obvious identifiers, savedViewId is optional). With only 3 params and this clarity, a 3 is fair – the description adds no param details, but none are critically needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the purpose: 'Read the implementation project end to end' and lists the specific components it retrieves (phases, next action, task list, sign-off rows, decision log, parallel-run status, and computed statuses). It distinguishes itself from sibling tools like implementation_project_list (lists projects) and implementation_project_create (creates).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says 'Gates on manage_implementation', which implies a prerequisite or dependency but is cryptic. It doesn't explicitly state when to use vs alternatives, but the tool is self-evidently a read endpoint for a specific project, distinct from the list sibling. A clearer 'use this to fetch a single project's full state' would push it to 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
implementation_project_listARead-only
List this workspace’s implementation projects as roster metadata (phase, first blocker + owner kind, days to cutover), blocked-first then by cutover proximity. Workspace-scoped: the cross-client roster composes N of these over A23 list_workspaces client-side. §H-TENANT: never reads across workspaces. Gates on read_master_data so a mandate member can compose the roster.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the readOnlyHint annotation by disclosing workspace isolation, the read_master_data permission gate, and the ordering/sorting behavior. This gives an agent a clear model of side effects and authorization requirements with no contradiction against annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the first sentence carries the action, resource, output fields, and ordering; subsequent sentences add scope, auth, and permission context. It uses domain jargon (§H-TENANT, A23, mandate) that may cost clarity for unfamiliar agents, so it is not a perfect 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with no output schema, the description covers the returned fields, ordering, workspace scope, cross-client composition, and permission gate. The main gaps are the undocumented status and savedViewId parameters and the lack of any pagination or return-envelope details. Overall strong, but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers 0% of parameters in descriptions, and the description only implies workspaceId through 'this workspace's' and 'Workspace-scoped'. The optional status and savedViewId parameters are never explained, so an agent cannot know how to filter the list or apply a saved view. Since schema coverage is low, the description needed to compensate and did not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb ('List'), names the exact resource ('this workspace's implementation projects'), and states the concrete shape of the result ('roster metadata: phase, first blocker + owner kind, days to cutover') plus the ordering ('blocked-first then by cutover proximity'). This makes it immediately distinguishable from implementation_project_get and the many other list tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit scoping guidance: it is workspace-scoped, 'never reads across workspaces', and explains that the cross-client roster is composed client-side over A23 list_workspaces. It does not explicitly contrast with implementation_project_get or other implementation siblings, but the workspace-isolation rule and composition pattern tell an agent when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
implementation_runbook_instantiateA
Instantiate a shipped runbook template into owned, dated, evidenced tasks: each task carries an owner kind (human|agent|system), a due date offset from the cutover date, a prerequisite, the evidence it requires and a contingency. The go/no-go and rollback tasks are undeletable (only not_applicable with a reason). The 180-day Umsatzabstimmung, the 240-day Berichtigung (Art. 72 MWSTG) and the prior-year Umsatzabstimmung land with computed statutory dates. Gates on manage_implementation.
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes | ||
| templateId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full behavioral burden. It discloses important behavior: tasks carry owner kind, evidence, prerequisites, contingencies, undeletable go/no-go and rollback tasks, and statutory date computation for specific Swiss tasks. It omits idempotency-replay behavior, but otherwise provides substantial behavioral detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the main action and then layers in concrete constraints. It is a long single description, but every clause adds meaningful behavioral detail rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, lack of annotations, lack of output schema, and four required parameters, the description covers the main effect, task constraints, statutory dates, and permission gating. A small gap remains around what the operation returns and how the idempotencyKey behaves on reuse.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the four parameters, but it does not explain workspaceId, projectId, templateId, idempotencyKey, or their semantics. The parameter names are somewhat self-explanatory, and 'templateId' is only indirectly referenced, but the important idempotencyKey behavior is not addressed at all.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Instantiate a shipped runbook template into owned, dated, evidenced tasks'. It enumerates the resulting task semantics and clearly differentiates this tool from generic task creators like tasks_create and implementation_* siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context by limiting the tool to 'shipped runbook templates' and mentioning the authorization gate, 'Gates on manage_implementation'. It does not explicitly name alternatives or state when-not conditions, but the context is sufficient to place the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
implementation_signoff_recordA
Record one human sign-off (append-only). kind is the FIXED enum: conversion_date, mwst_method, mapping_approval, tieout, contact_merge, go_nogo, rollback_trigger, parallel_run_close, source_cancellation. EVERY kind is a human act: an agent (or the system seat) is REFUSED (signoff_needs_human), never staged (P8). The tieout and parallel_run_close kinds bind the current reconciliation evidence hash, so a later check that changes the evidence voids the sign-off. Gates on commit_migration (signing is the human half of committing).
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | ||
| projectId | Yes | ||
| evidenceRef | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With zero annotations, the description carries the full burden and delivers richly: append-only semantics, agent refusal with signoff_needs_human, never staged (P8), the evidence-binding behavior that voids on later change for tieout/parallel_run_close kinds, and the commit_migration gate. This is far beyond what the bare schema provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four dense sentences with the core purpose front-loaded ('Record one human sign-off (append-only)') before the enum list and behavioral constraints. Every sentence adds distinct information; there is no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Strong for a 5-parameter tool with no output schema and no annotations: enum, refusal, evidence-voiding, and workflow gate are all covered. Gaps remain around idempotencyKey semantics and what a successful call returns, which an agent invoking an append-only record would need to know.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It fully defines the kind enum and gives evidenceRef real meaning via the evidence-hash binding. However, idempotencyKey — critical for an append-only idempotent record — is never explained, and workspaceId/projectId scoping is only implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Record one human sign-off (append-only).' The fixed enum of nine kinds precisely scopes what acts qualify, and the append-only trait separates it from mutable record or decision tools. It is clearly distinguishable from sibling tools like implementation_decision_record or migration_record_approval.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides strong contextual usage signals: it is the 'human half of committing' and gates on commit_migration, and the refusal note ('an agent (or the system seat) is REFUSED') tells the agent this tool is for recording human acts, not for agent-initiated staging. However, it does not explicitly name alternative tools or state when a sibling (e.g., implementation_decision_record) should be chosen instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
implementation_task_setA
Create or update one implementation task (pass taskId to update). fields carries phase, title, ownerKind, ownerRef, due, status, evidenceKind, evidenceRef, contingency and reason. Completing a task whose prerequisite is still open refuses (prerequisite_open, naming it); completing one that requires evidence without it refuses (evidence_required, naming the kind); marking a task not_applicable needs a recorded reason (not_applicable_needs_reason). Gates on manage_implementation.
| Name | Required | Description | Default |
|---|---|---|---|
| fields | No | ||
| taskId | No | ||
| projectId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it discloses that completion is refused when prerequisites are open or evidence is missing, that not_applicable requires a reason, and that manage_implementation permission is required. It does not mention idempotency semantics or side effects beyond the task object itself, but the disclosed validation behavior is substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the core create/update behavior, followed by failure rules and the permission gate. Each sentence contributes meaningful information, though the middle sentence listing fields and the later constraint sentence are slightly compressed and could be clearer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of annotations and output schema, the description covers the core operation, field surface, key validation rules, and authorization gate. It is missing some context around idempotencyKey and field value semantics, but an agent has enough to call the tool correctly for common create/update flows.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does list the fields inside the fields object and note that taskId selects update mode. However, it adds little meaning for required parameters like workspaceId, projectId, and idempotencyKey, and the listed field names (ownerKind, evidenceRef, contingency) remain semantically thin.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb-resource pair, 'Create or update one implementation task', and clarifies the difference between the two modes with 'pass taskId to update'. This clearly distinguishes it from generic task tools like tasks_create/tasks_update by the 'implementation' scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies its usage context by naming implementation tasks and the manage_implementation gate, but it never explicitly states when to choose this tool over alternatives such as tasks_create, tasks_update, or implementation_project_* tools. No when-not-to-use guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_camtA
Import a camt.053 (statement) or camt.054 (debit/credit notification) file onto a registered Bankkonto (A19): parses the XML, checks the file's IBAN against the account, and persists every booked entry, each keyed on its own bank-assigned identity (AcctSvcrRef, falling back to NtryRef, falling back to a content hash) so an overlapping statement or the camt.054-then-camt.053 pair cannot re-import the same booking twice. A byte-identical re-import of the same statement (Stmt/Id + ElctrncSeqNb + page number) is a safe no-op (duplicate:true); a same-identity statement whose content has genuinely changed refuses with statement_amended, naming what changed. A genuine same-day twin the identity key would otherwise treat as a repeat is admitted only via allowDuplicateEntries:true. A batch entry (several TxDtls in one booking) fans out into one txn per TxDtls. A booked, non-reversal CREDIT is routed into A21's QR-Abgleich queue automatically; everything else waits for suggest_matches / confirm_match / create_entry_for_txn. Any entry this parser could not read, or an entry it skipped as a duplicate, is named in the skipped[] list, never silently dropped.
| Name | Required | Description | Default |
|---|---|---|---|
| xml | Yes | ||
| workspaceId | Yes | ||
| bankAccountId | Yes | ||
| idempotencyKey | Yes | ||
| allowDuplicateEntries | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With zero annotations, the description carries the full burden and does so excellently: it discloses the dedup identity chain (AcctSvcrRef→NtryRef→content hash), no-op re-import behavior (duplicate:true), the statement_amended refusal, the allowDuplicateEntries escape hatch, batch entry fan-out, automatic QR-Abgleich routing for credits, and the skipped[] list with no silent drops.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a dense single paragraph, but every clause earns its place for a tool this complex. It is front-loaded with the primary purpose before diving into dedup, error modes, and routing. Slight restructure into bullets could aid scanning, but there is no padding or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with no annotations and no output schema, the description covers dedup semantics, error modes, fan-out, routing, and failure handling thoroughly. It names return flags (duplicate:true, statement_amended, skipped[]) but never describes the full success response shape, and the idempotencyKey behavior is not fully explained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It richly explains xml (camt.053/054 content), bankAccountId (registered Bankkonto A19), and allowDuplicateEntries (genuine same-day twin admission). However, idempotencyKey's role is not explicitly tied to the dedup/re-import logic, and workspaceId is left to its name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource+scope: importing camt.053/camt.054 XML onto a registered Bankkonto, parsing it, checking IBAN, and persisting booked entries. It is sharply distinguished from siblings like list_bank_statements, export_statement, and bank_sync by describing the ingestion+dereplication role rather than listing/exporting/syncing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description names the downstream workflow explicitly (suggest_matches / confirm_match / create_entry_for_txn) and explains when re-imports are safe no-ops vs when they are refused. It does not give an explicit 'use X instead' exclusion, but the context for when this tool is the right entry point is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_exchange_ratesA
Import a published BAZG rate payload into the rate store as source='rate_api'. payload is the response body from the endpoint describe_rate_feed names, verbatim. The payload identifies its own series; pass series to have a mis-paired payload refused rather than mis-dated. Daily rates are stored on each date in the payload's validity window (not on the determination date), monthly averages once on the first of their month, and per-unit quotations (100, 1000, 10000) are scaled. A published rate TILL cannot hold exactly at twelve decimal places is reported rather than rounded; the whole BAZG daily series is within that (its deepest is nine places).
| Name | Required | Description | Default |
|---|---|---|---|
| series | No | ||
| payload | Yes | ||
| currencies | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral disclosure and delivers richly. It details how daily rates are dated (validity window, not determination date), monthly averages (first of month), per-unit scaling (100/1000/10000), and the precision behavior (reported rather than rounded). It also discloses a refusal path for mis-paired payloads when `series` is supplied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded, stating the purpose first and then detailing parameter semantics and storage rules. Each sentence contributes new information, but the final sentence contains a slightly garbled phrase ('A published rate TILL cannot hold...') that marrs an otherwise tight structure. It remains concise for the amount of behavioral detail packed in.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of output schema and annotations, the description covers the key operational details: payload provenance, series validation, storage rules, scaling, and precision handling. It does not explain the `currencies` parameter or the return value/confirmation behavior, which would round out a fully complete picture. Overall, it is thorough enough for an agent to invoke correctly with minimal guesswork.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description compensates well for the two most complex parameters: `payload` is defined as the exact response body from `describe_rate_feed`, and `series` is explained as a validation guard. The remaining parameters (`currencies`, `workspaceId`, `idempotencyKey`) are not elaborated, but they are either common conventions or optional; the description adds meaningful semantics where ambiguity is highest.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pairing: 'Import a published BAZG rate payload into the rate store as `source='rate_api'`.' This clearly distinguishes it from sibling rate tools like `record_exchange_rate` or `describe_rate_feed` by specifying the exact input and storage action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The tool defines its primary use case: the `payload` is the verbatim response from `describe_rate_feed`'s named endpoint, making it clear when this tool is appropriate. It also explains the optional `series` parameter as a safety check for mis-paired payloads. It stops short of explicit exclusions or naming alternative tools, but the context is strong enough for an agent to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_opening_balancesA
Import already-parsed migration rows as the opening position, posting exactly what preview_opening_import showed. Takes rows as an array of column-value objects, NOT a CSV string: this verb parses no file, and turning an export into rows is the caller's step. Refuses with unmapped_account when any row names an account the chart does not have, because importing the rest would post a position that balances only because a row was dropped. A row with no account at all (what a subtotal or Total line looks like) is refused as invalid_row naming that row, rather than blamed on the mapping. format accepts csv or bexio, which select default column names; mapping overrides them field by field for an export with different headers. Retrying the same idempotencyKey returns the existing entry and never a second opening position.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| rows | Yes | ||
| dryRun | No | ||
| format | No | ||
| mapping | No | ||
| reference | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| differenceAccount | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully covers behavioral aspects: it discloses refusal conditions (unmapped_account, invalid_row), explains format and mapping overrides, and details idempotency behavior (retry returns existing entry). This is comprehensive and goes beyond minimal expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence adds value, covering purpose, input format, error cases, format/mapping, and idempotency. It is logically organized and front-loaded with the core purpose. Slightly verbose but justified given the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 parameters, nested objects, and no output schema, the description is fairly complete. It explains main behavior, error handling, format, and idempotency, but lacks details on optional parameters like dryRun and differenceAccount, and does not describe return values. Still sufficient for basic correct usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage, so the description must explain parameters. It explains rows (array of objects), format (csv or bexio), mapping (overrides), and idempotencyKey, but omits asOf, dryRun, reference, and differenceAccount. While key parameters are covered, several remain unexplained, leaving gaps for the agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool imports already-parsed migration rows as the opening position, referencing the preview tool. It distinguishes itself by explicitly noting it does not parse CSV files, which differentiates it from other import tools. The verb 'import' and resource 'opening position' are specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage after preview_opening_import and clarifies the caller must convert exports to rows, but it does not explicitly name alternative tools or state when not to use it. The reference to preview_opening_import provides a clear workflow, but explicit exclusions would improve it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_open_itemsA
Import open items (ar or ap) for a migration plan: each already-mapped row becomes an origin=migrated document (AR, at issued) or vendor bill (AP, at posted) that posts NOTHING of its own (posted_entry_id/entry_id NULL). Its only ledger effect is A04 opening 1100/2000 line; the migrated open total must tie to that line to the Rappen (a nonzero difference is a red control, never a plug posting). Atomic and idempotent: a single refused row imports nothing, and a re-run with the same idempotencyKey replays the original outcome (no duplicate rows). Rides commit_migration; denylisted from automation. priorYearDetail live carries an already-settled prior-year row as a settled migrated item that nets to zero open; archive (default) refuses it toward the G13 archive. CONSEQUENCE: Writes the migrated open AR and AP items into the live opening balance; the batch cannot be un-imported.
| Name | Required | Description | Default |
|---|---|---|---|
| rows | Yes | ||
| side | Yes | ||
| planId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| priorYearDetail | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses extensive behavioral details: posts nothing of its own, the only ledger effect is A04 opening line, atomic and idempotent (single refused row imports nothing, re-run with same idempotencyKey replays original outcome), denylisted from automation, priorYearDetail behavior, and the CONSEQUENCE that writes into live opening balance and cannot be un-imported. This is exemplary transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and long, but every clause earns its place given the tool's complexity. It is front-loaded with purpose, then details behavior, constraints, and consequences. While lengthy, it is efficient and well-organized, not verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is complex with migration, idempotency, side effects, and prior-year handling. The description covers atomicity, idempotency, side effects, control totals, and the irreversible consequence. It lacks details on the exact return value or allowed side values, but given no output schema and the depth provided, it is fairly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains idempotencyKey (replays original outcome), priorYearDetail (settled vs archive refusal), rows (each already-mapped row), and side (ar or ap). It does not explicitly describe planId and workspaceId, but these are common IDs that are self-explanatory in context. The description adds meaning beyond the schema, though not exhaustive for all six parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Import open items (ar or ap) for a migration plan' and details exactly what happens to each row (origin=migrated document or vendor bill, posts nothing of its own, A04 opening line). This is specific and distinct from sibling tools, which are also migration/import related but perform different steps. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says 'for a migration plan' and notes it 'Rides commit_migration; denylisted from automation', providing clear context on when to use it. However, it does not explicitly compare to alternatives or state when not to use it, so it lacks explicit exclusions. Still, the context is strong enough to route an agent correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
income_statementARead-only
Erfolgsrechnung (income statement) for a period, in the OR Art. 959b Abs. 2 Gesamtkostenverfahren layout: the ten account-backed positions in their prescribed order, each under its statutory wording, plus reingewinnMinor, which is position 11, Jahresgewinn oder Jahresverlust. Unlike the Bilanz this Absatz is a flat list with no sub-level, so all eleven positions are present and in the prescribed order. Two things to say about the eleventh rather than claim: it is the COMPUTED result and not a section of its own, so no account falls into it, and the exported Erfolgsrechnung heads it with the ENACTED wording resolved by sign, Jahresgewinn above zero and Jahresverlust below it, keeping the full Jahresgewinn oder Jahresverlust verbatim at exactly zero. Each single word is the enacted wording with the branch the book did not take dropped, because the oder enumerates the two outcomes a result can have rather than naming one heading. The conventional Treuhand wording Reingewinn oder Reinverlust is not printed. The Bilanz carries the same figure under the same enacted wording as an Eigenkapital position, so the two statements of one Jahresrechnung give one number one name. A RESIDUAL section, uebrige_positionen, follows the ten and holds anything that fits none of them, so no account is dropped. It is the "weitere Positionen" of OR Art. 959b Abs. 5, which is a different Absatz and not one of the eleven: it is neither position 11 nor a twelfth prescribed position. Every position is signed as its CONTRIBUTION TO PROFIT (revenue positive, expense negative, and the three positions the statute writes as Aufwand AND Ertrag shown net), so a loss is a negative figure and never an absolute value. This is the verb that answers "wie war der Gewinn in Q2?". Pass compareTo={periodStart,periodEnd} for the prior-period column. The Absatzerfolgsrechnung of OR Art. 959b Abs. 3 is not built, and this is the filed statement, so it accepts no alternate grouping.
| Name | Required | Description | Default |
|---|---|---|---|
| compareTo | No | ||
| periodEnd | Yes | ||
| periodStart | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide readOnlyHint:true. The description adds substantial behavioral context beyond that: the exact layout, the eleven positions in prescribed order, that reingewinnMinor is a computed result and not an account section, sign convention (contribution to profit, negative for losses), the handling of the zero case with the full enacted wording, and that no account is dropped due to the residual section. It also explains why one wording is dropped under the oder enumeration. This is rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place for a complex statutory reporting tool. It is front-loaded with the core purpose and layout, then explains the computed eleventh position, sign convention, and residual section. It could be tightened slightly — some statutory nuance about the weniger/oder wording is repeated across a couple of sentences — but it is structured and not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the tool — a statutory Swiss income statement with an output schema absent and schema parameter coverage at 0% — the description is remarkably complete. It covers the positions, order, sign convention, result handling, residual section, comparison parameter, and exclusions. An agent could invoke it correctly and interpret its output without needing additional documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden for parameter semantics. It explains the meaning of periodStart/periodEnd by saying it answers 'wie war der Gewinn in Q2?' and explains that compareTo={periodStart,periodEnd} selects the prior-period column. However, it does not describe workspaceId, though that is likely obvious from context. It compensates well for the schema gap on the temporal and comparison parameters but leaves workspaceId implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a specific verb+resource — the income statement (Erfolgsrechnung) for a period — and goes well beyond the title by naming the exact statutory layout (OR Art. 959b Abs. 2 Gesamtkostenverfahren), the eleven positions, and the residual section. It also distinguishes itself from sibling tools: it says this is the filed statement and that the Absatzerfolgsrechnung of OR Art. 959b Abs. 3 is not built, and contrasts with Bilanz. An agent can select it confidently.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use it — to answer "wie war der Gewinn in Q2?" — and how to request the prior-period column via compareTo={periodStart,periodEnd}. It also names what it is not: the Absatzerfolgsrechnung is not built, it accepts no alternate grouping, and the Treuhand wording is not printed. Siblings like balance_sheet, trial_balance, general_ledger are implicitly distinguished by the explicit scope and layout.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
install_pluginA
Installiere eine Erweiterung (US-G02.1): check the pinned sha256 against the bundle (manifest_checksum_failed on a mismatch, nothing persisted), REJECT any capability naming a reserved money-path tool (capability_forbidden, no_second_posting_path, the WHOLE install fails, P3) or colliding with an existing tool/screen/source name (capability_name_conflict), store permissions.granted as the INTERSECTION of grantedScopes and the manifest`s requested scopes (never a superset), resolve compat_range against the core version, and persist a plugin_manifests row (status installed, or incompatible when the range fails). A newer version of an already-installed plugin supersedes the row in place and writes an audit_log entry before it changes. CONSEQUENCE: Installs third-party code into the workspace with the scopes it was granted.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | ||
| packageRef | Yes | ||
| workspaceId | Yes | ||
| grantedScopes | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and delivers exceptionally: it discloses the sha256 verification with manifest_checksum_failed and 'nothing persisted', rejection rules with specific error codes (capability_forbidden, no_second_posting_path, capability_name_conflict), the scope-intersection security constraint ('never a superset'), compatibility resolution, row persistence semantics, superseding behavior with audit logging, and a plain-language CONSEQUENCE statement about installing third-party code. This far exceeds what annotations would typically provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a dense run-on wall of text — roughly 150 words in one meandering sentence packed with CAPS emphasis and parenthetical error tokens like (manifest_checksum_failed on a mismatch, nothing persisted). Every clause carries behavioral value and it is front-loaded with the purpose, but the lack of sentence breaks and mixed-language opening make it difficult to scan efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a security-sensitive tool with no annotations, no output schema, and 5 undocumented parameters, the description covers an impressive amount: failure modes, concurrency/superseding behavior, scope safety, persistence, and audit trail. The notable gaps are idempotencyKey semantics (critical for a state-changing install), the return value of a successful install, and the meaning of source. Strong but not fully complete for this complexity level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it partially does: grantedScopes gets real meaning via the intersection rule with the manifest's requested scopes, and packageRef is hinted at through the pinned-sha256 bundle check. However, idempotencyKey — a required parameter — is never mentioned, and source and workspaceId receive no semantic explanation. Meaningful for two of five parameters, negligible for the rest.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Installiere eine Erweiterung' (install an extension), a specific verb+resource, and then details the install behavior (persisting a plugin_manifests row), which clearly separates it from siblings like preview_plugin_install, enable_plugin, and uninstall_plugin. It doesn't explicitly name a sibling to differentiate against, but the level of operational detail makes the distinction unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance is given. The description never mentions safer alternatives like preview_plugin_install for dry-run validation, refresh_plugin_compat for non-committal compatibility checks, or uninstall_plugin for reversal. The only implied usage context is that this is the actual install operation, with no exclusions or routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_adjustA
Post a single reason-coded quantity adjustment. Supply itemId, locationId, qtyDelta (signed integer thousandths; positive raises on-hand, negative lowers it), reasonCodeId (mandatory, must be active), optional note (required when the reason has requiresNote), optional lotId / serialId, optional unitCostMinor (a Rappen snapshot for the J03 valuation layer), optional effectiveDate (ISO date, default today). It validates the reason and note, asserts the period is open, then mints ONE J02 inventory_move (movement_type adjustment) and records an inventory_adjustment row linking the movement to the reason. Rejections mint nothing: reason_inactive / not_found / no_active_reasons (bad or missing reason), note_required, invalid_qty (qtyDelta 0), insufficient_stock (would overdraw with allow_negative_stock off), period_locked. Idempotent under idempotencyKey (a replay returns the original, posting no second movement).
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| lotId | No | ||
| itemId | Yes | ||
| qtyDelta | Yes | ||
| serialId | No | ||
| locationId | Yes | ||
| workspaceId | Yes | ||
| reasonCodeId | Yes | ||
| effectiveDate | No | ||
| unitCostMinor | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden and exceeds it. It discloses side effects precisely (mints ONE J02 inventory_move with movement_type adjustment, records an inventory_adjustment row), atomic failure behavior ('Rejections mint nothing' with all six rejection reasons), and idempotency semantics (replay returns original, posts no second movement). This is model-grade transparency for a mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The text is appropriately sized for an 11-parameter mutation with rich semantics and is front-loaded with the core action and parameter semantics. Every sentence earns its place: parameter meanings, validation rules, side effects, rejection list, and idempotency behavior. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the high complexity (11 params, 6 required, no output schema, no annotations), the description covers parameter semantics, validation rules, all failure modes, side effects, and idempotency — everything an agent needs to invoke it correctly. The only minor gap is the general return payload shape, but the exact rows created are disclosed, which effectively answers what the result is. No output schema exists to offload this, yet the description is still complete enough to act on.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description is the only source of parameter meaning, and it compensates fully: qtyDelta is specified as signed integer thousandths with direction semantics, reasonCodeId's activeness constraint, note's conditional requirement, unitCostMinor's valuation-layer context (J03 Rappen snapshot), and effectiveDate's ISO format and default. Only workspaceId and idempotencyKey are left to the schema property names, which are self-explanatory. This far exceeds the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise action ('Post a single reason-coded quantity adjustment'), names the resource, and disambiguates scope by stressing 'single' — contrasting with the batch sibling inventory_adjust_batch and the reversal sibling inventory_adjust_reverse. The behavioral specification of exactly which rows are minted (ONE J02 inventory_move + inventory_adjustment row) makes the tool's role unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides rich usage context: it specifies prerequisites (reason must be active, period must be open, note required when requiresNote), states the 'single' vs batch scope, and enumerates the valid conditions for success and all rejection reasons. However, it never explicitly names an alternative tool or states 'use X instead when ...', so the differentiation from siblings like inventory_move or inventory_adjust_reverse is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_adjust_analysisARead-only
Aggregate adjustment quantity and value impact over a date window, grouped by any of reason, category, item, location (default reason). Returns per-group adjustmentCount, net qtyDelta and valueImpactMinor (SUM(qtyDelta * unitCostMinor) in Rappen). Reversal rows are included so a reversed adjustment nets to zero, giving an honest shrinkage / write-down figure for Swiss Inventar review.
| Name | Required | Description | Default |
|---|---|---|---|
| toDate | No | ||
| groupBy | No | ||
| fromDate | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses important behavior: reversal rows are included and net to zero, valueImpactMinor uses Rappen, and the default grouping is reason. It also reveals that the metric is 'honest shrinkage / write-down,' which helps the agent reason about what the numbers mean.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three dense sentences with no filler. Each clause adds unique value: aggregation scope, grouping options, return fields, formula, unit, and reversal behavior. It is efficiently front-loaded and stays appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description appropriately explains return metrics and the reversal semantics, which is the core context an agent needs. It is slightly incomplete because it does not clarify date-range bounds, the required workspaceId role, or groupBy array syntax, but the main behavioral contract is well covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain the groupBy values (reason, category, item, location) and mentions a date window implicitly mapping to toDate/fromDate, but it never names workspaceId or describes expected formats, date inclusivity, or how the groupBy array should be structured. Partial compensation, not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb ('Aggregate') and clearly names the resource and output: adjustment quantity and value impact over a date window, grouped by reason/category/item/location. It also defines the return metrics and even the exact formula, making the tool's purpose unambiguous and distinguishable from raw adjustment-list and adjustment-write siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context: it is the tool for an aggregate, per-group summary of adjustment impact and is framed for 'Swiss Inventar review,' which implies when a summarized figure is needed. It does not explicitly name alternatives or state when not to use it, but the aggregate-vs-list distinction is clear enough from the wording.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_adjust_batchA
Post a multi-line adjustment atomically under one generated batchId. lines[] each carry itemId, locationId, qtyDelta, reasonCodeId (may differ per line) and optional lotId / serialId / note / unitCostMinor. Every line is validated (reason active, note policy) BEFORE any mint and the period is asserted once; all lines mint in ONE transaction. If any line fails (inactive reason, missing required note, insufficient stock, invalid qty), the WHOLE batch rolls back and nothing is written. effectiveDate defaults to today. Idempotent under idempotencyKey.
| Name | Required | Description | Default |
|---|---|---|---|
| lines | Yes | ||
| description | No | ||
| workspaceId | Yes | ||
| effectiveDate | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well. It discloses atomicity, validation-before-mint, period assertion, rollback on failure, default effectiveDate, and idempotency under idempotencyKey. These are critical behavioral traits that an agent must know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-organized paragraph that front-loads the purpose and then provides essential behavioral details. Every sentence adds value; there is no redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a batch atomic operation, the description covers the core behaviors (validation, rollback, idempotency) and parameter semantics. It does not mention the response format (e.g., that it returns the generated batchId), but that is a minor omission given the richness of other context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does. It explains the lines array structure (itemId, locationId, qtyDelta, reasonCodeId, optional lotId/serialId/note/unitCostMinor), effectiveDate default, and idempotencyKey purpose. It covers most parameters meaningfully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Post a multi-line adjustment atomically under one generated batchId'), which distinguishes it from single-line inventory_adjust and other inventory tools. It specifies the resource (inventory adjustment batch) and the key characteristics (multi-line, atomic).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for multi-line adjustments and mentions atomicity and idempotency, but does not explicitly name alternatives or state when not to use it. It lacks a clear 'use X instead when you have a single line' type of guidance, which would differentiate it from inventory_adjust.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_adjust_listARead-only
List manual adjustments (newest first) with reason code + name + category, item and location names, the minted movement id, note, signed qtyDelta and any reversal back-link. Filterable by fromDate / toDate (effective date), itemId, locationId, reasonCodeId, category and batchId. Reversal rows are excluded unless includeReversals. Paginated by limit (default 100, max 500) and offset.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| itemId | No | ||
| offset | No | ||
| toDate | No | ||
| batchId | No | ||
| category | No | ||
| fromDate | No | ||
| locationId | No | ||
| workspaceId | Yes | ||
| reasonCodeId | No | ||
| includeReversals | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, which is consistent. The description adds behavioral details beyond annotations: reversal rows are excluded unless includeReversals is set, and pagination with default limit 100 and max 500. It also mentions sorting (newest first). These are valuable behavioral disclosures that help an agent anticipate results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph that front-loads the purpose, then lists returned fields, filters, and pagination. It is informative without being overly verbose, though it could be slightly more structured (e.g., bullet points). Overall, it earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with 11 parameters, 1 required, and no output schema, the description is remarkably complete. It specifies return fields, filters, sorting, pagination limits, and the reversal exclusion behavior. An agent has enough information to call it correctly and interpret results. The only minor omission is the exact response structure, but the listed fields suffice.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It does so by listing most parameters and their meanings: fromDate/toDate as effective date, itemId, locationId, reasonCodeId, category, batchId, and pagination params with defaults. It does not detail workspaceId but that is required and obvious. The description adds semantic meaning beyond the bare schema, though not exhaustive for every parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action: listing manual adjustments, sorted newest first, with explicit fields returned (reason code, item/location names, movement id, qtyDelta, etc.). It differentiates from siblings like inventory_adjust (which likely creates) by specifying 'manual adjustments' and listing fields, making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on what the tool does and its filtering capabilities, including date range, item, location, reason, category, and batchId. It does not explicitly name alternatives or state when not to use it, but the purpose is evident and the filtering options give a good sense of use cases. It also notes reversal exclusion, adding usage nuance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_adjust_reverseA
Reverse a single adjustment (adjustmentId) or a whole batch (batchId). Each reversal mints an opposite-sign J02 movement carrying its OWN reasonCodeId (mandatory, active) and records a new inventory_adjustment row with reverses_adjustment_id set; the original row and movement are never touched. Reversing a single adjustmentId that belongs to a batch reverses just that line; reversing by batchId reverses every not-yet-reversed line. Already reversed is already_reversed. Batch reversal is atomic. effectiveDate defaults to today (a reversal posts a correction in the current open period). Idempotent under idempotencyKey.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| batchId | No | ||
| workspaceId | Yes | ||
| adjustmentId | No | ||
| reasonCodeId | Yes | ||
| effectiveDate | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly states side effects: 'Each reversal mints an opposite-sign J02 movement... records a new inventory_adjustment row... the original row and movement are never touched.' It also discloses atomicity, idempotency, and the default effectiveDate. This is exceptionally transparent for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph that front-loads the core purpose and then details behavior without redundancy. Every sentence adds value: the two modes, the mechanics of reversal, atomicity, default date, and idempotency. It is efficient and well-structured despite its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the operation and the absence of annotations and output schema, the description covers all essential aspects: which identifiers to use, what happens internally, atomicity, idempotency, and the effectiveDate default. It does not describe return values (no output schema exists), but that is not required. The description is complete for an agent to correctly invoke the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the critical parameters: adjustmentId vs batchId, reasonCodeId (mandatory, active), effectiveDate (defaults to today), and idempotencyKey (idempotent). However, it does not explain workspaceId (though required, likely standard) or note (optional). The description adds significant meaning beyond the schema for most parameters, meriting a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reverses an inventory adjustment, specifying two distinct modes: by adjustmentId or by batchId. It distinguishes this from inventory_adjust and inventory_adjust_batch (which create adjustments) by explicitly stating 'Reverse a single adjustment or a whole batch.' The verb is precise and the resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use guidance for both parameters: 'Reversing a single adjustmentId that belongs to a batch reverses just that line; reversing by batchId reverses every not-yet-reversed line.' It also notes atomicity for batch reversal and that 'Already reversed is already_reversed' implying you cannot reverse again. This leaves no ambiguity about parameter selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_alertsARead-only
Der einheitliche Alarm-Feed (unified alerts, US-J07.9): one tool an agent can poll for everything that currently needs attention. It merges low-stock, anomalies, overdue stocktakes, valuation drift and near-expiry lots into a single de-duplicated, severity-prioritised list, each alert carrying a stable alert_key (for optional future acknowledgement), type, severity, title, summary, entity refs, detected_at and a suggested_action hint (create_requisition, run_valuation, complete_stocktake, ...). Optional severity / types / limit filters. Pure computation; no persistent alert store in Phase 1. Empty is an ok empty list. Posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| types | No | ||
| severity | No | ||
| thresholds | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is exceptionally transparent: 'Pure computation; no persistent alert store in Phase 1. Empty is an ok empty list. Posts nothing.' This discloses the read-only, stateless nature beyond the readOnlyHint annotation. It also notes the optional future acknowledgement feature (stable alert_key), setting expectations about current behavior. This fully aligns with the annotation, adding context about persistence and side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then details the output structure, filters, and behavioral notes. Each sentence adds value, but it is somewhat long and could be tightened. For example, the list of alert types could be shortened, and the parenthetical spec reference '(US-J07.9)' is non-essential. Overall, it is well-organized but slightly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema and parameter descriptions, the description carries a heavy burden. It covers the return structure (alert fields), filter options, and behavioral guarantees (read-only, no persistence, empty list). It does not explain the 'thresholds' parameter, which is a notable omission for a tool that aggregates alerts. It also doesn't provide examples or enumerate valid severity/type values, but these might be inferred from context. Overall, it is comprehensive but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% coverage for parameter descriptions, so the description must compensate. It mentions 'Optional severity / types / limit filters', which covers three of the five parameters (limit, types, severity), but omits 'thresholds' and the required 'workspaceId'. While it implies these filters exist, it does not explain their values or formats. The description adds some meaning but leaves gaps, especially for thresholds, which could be crucial for tuning alerts.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: a unified alert feed that merges multiple alert types into a single, de-duplicated, severity-prioritized list. It names the specific alert categories (low-stock, anomalies, overdue stocktakes, valuation drift, near-expiry lots) and lists the output fields. This distinguishes it from sibling tools like detect_anomalies or inventory_anomalies, which are more specialized. The verb 'poll' and resource 'everything that currently needs attention' make it unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says it is 'one tool an agent can poll for everything that currently needs attention', which is a clear when-to-use statement. It also mentions optional filters (severity, types, limit) to refine results. However, it does not explicitly name alternative tools or conditions when NOT to use it, such as when a specific anomaly detection is needed instead. The guidance is implicit rather than comparative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_anomaliesARead-only
Die Anomalieliste (anomalies, US-J07.5): a prioritised list of DERIVED inventory exception events since a look-back date (default 30 days), each with type, severity (info | warning | critical), title, summary, entity refs, detected_at and a structured payload. Detected types (Phase 1, all derived, no extra table): negative_stock (a live on-hand below zero), large_issue / large_adjustment (absolute qty or value above threshold), unlinked_high_value_issue (a high-value issue with no source document), valuation_drift (cross-ref valuation status), open_stocktake_overdue (a J04 session open past its window) and lot_near_expiry (an open lot inside the J01 expiry window). Thresholds are built-in defaults, overridable per call, never persisted. Empty is an ok empty list. Deterministic for a given snapshot; posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| since | No | ||
| types | No | ||
| severity | No | ||
| thresholds | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses determinism for a given snapshot, absence of side effects ('posts nothing'), non-persistence of thresholds ('never persisted'), and that an empty result is valid. It also explains that thresholds are built-in defaults overridable per call, adding useful behavioral context not available from annotations or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-dense with no filler, front-loading the core 'prioritised list of DERIVED inventory exception events. It is somewhat long and opens with a German phrase that adds little for a non-German agent, but every subsequent clause contributes meaningful detail about types, behavior, and output shape.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without an output schema, the description covers the result fields, the full type catalogue, threshold behavior, determinism, and side-effect safety. It leaves limit and workspaceId semantics implicit, but the six parameters are simple enough that an agent can invoke the tool correctly with the provided detail and schema names.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the parameter-documentation burden and mostly meets it: it explains the since look-back default, enumerates the types filter values, names severity values, and clarifies thresholds are per-call overrides. It does not document limit or workspaceId semantics explicitly, but their purpose is inferable from their names, so the gap is minor.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it returns a prioritized list of DERIVED inventory exception events since a look-back date, and enumerates the exact detected anomaly types with their meanings. It differentiates from likely siblings like detect_anomalies by asserting 'Phase 1, all derived, no extra table' and 'posts nothing', making its role as a read-only anomaly list unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives such as detect_anomalies, procurement_anomalies, inventory_alerts, or attention_summary. The description implies it is for reading derived inventory exceptions, but it never states why this list should be chosen over a sibling, which is a significant gap in a sibling set of this size.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_available_serialsARead-only
The serials whose current derived status is available (or a supplied status list), optionally filtered by itemId, current locationId or lotId. A pure read, side-effect free: the allocation/reservation query an agent runs before picking units.
| Name | Required | Description | Default |
|---|---|---|---|
| lotId | No | ||
| itemId | No | ||
| status | No | ||
| locationId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already signals safety; the description adds value by stating 'a pure read, side-effect free' and by explaining that the status is a 'current derived status' rather than a raw stored field. It also discloses its intended place in the picking workflow. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two efficient sentences: the first states the resource, default, and filters; the second establishes safety and use case. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is nearly complete for a read-only filtered list tool: it covers purpose, filters, default vs supplied status, and workflow context. Minor gaps remain because there is no output schema and no mention of return shape, pagination, or ordering, but these are not critical for an agent deciding whether to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% property description coverage, so the description carries the burden. It compensates by explaining status as a supplied status list, and identifying itemId, current locationId, and lotId as optional filters, including the important 'current' nuance for locationId. It does not detail status values or workspaceId, but most parameter meaning is conveyed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the precise resource (serials), their status condition (current derived status available or one of a supplied list), and the optional filters (itemId, current locationId, lotId). It is immediately distinguishable from sibling serial and inventory tools by framing it as the allocation/reservation read query.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear use context: 'the allocation/reservation query an agent runs before picking units.' It also clarifies optional filtering dimensions. It does not explicitly list sibling alternatives or say when not to use it, but the context is strong enough for an agent to select it correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_balanceARead-only
On-hand quantity as the pure SUM of matching movements (never a cached column). Filter by any subset of itemId, locationId, lotId, serialId, and an inclusive asOf date (effective_date <= asOf); omit asOf for now. Returns qtyOnHand for the whole filter plus a per item x location breakdown. A guessed foreign workspace id returns zero / empty, never data.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| lotId | No | ||
| itemId | No | ||
| serialId | No | ||
| locationId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true already present, the description adds genuine value: exact computation semantics (SUM of matching movements, never cached) and a specific failure behavior ('A guessed foreign workspace id returns zero / empty, never data'). Minor gaps remain (no note on output size, units, or pagination), but annotations reduce the burden and the added context is substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, each carrying load-bearing information: the computation semantics are front-loaded, filter rules are compressed into one clause, and the return shape plus workspace failure mode round it out. Zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only query tool with no output schema, the description nicely compensates by stating the return shape (whole-filter qty plus breakdown). Filtering, asOf inclusivity, and workspace scoping are all covered. The only gaps are minor (numeric format/units of qtyOnHand) and no explicit sibling routing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate — and it covers every parameter: the four filter dims (itemId, locationId, lotId, serialId), asOf (inclusive with effective_date <= asOf, omittable), and workspaceId (foreign id yields empty). Only value formats/examples are missing, which is minor.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('On-hand quantity as the pure SUM of matching movements') and clearly scopes what is returned (qtyOnHand for the filter plus an item x location breakdown). It differentiates from cached-report siblings by explicitly declaring 'never a cached column'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Filter semantics are well explained ('any subset... inclusive asOf (effective_date <= asOf); omit asOf for now'), and 'never a cached column' implies when to prefer it over cached alternatives. However, no sibling tool is named and no when-not-to-use guidance is given, which matters given several similar inventory-balance siblings (stock_on_hand, inventory_balance_by_location, inventory_stock_position).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_balance_by_locationARead-only
On-hand quantity broken down by location, plus a warehouse roll-up. A pure read model over the OP2 stock_movement ledger (SUM of signed qty per item x location), never a cached column. Filter by itemId, warehouseId or locationId; includeZero returns pairs that net to zero. Quantities are integer units.
| Name | Required | Description | Default |
|---|---|---|---|
| itemId | No | ||
| locationId | No | ||
| includeZero | No | ||
| warehouseId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses that this is a pure read model over the OP2 stock_movement ledger, that it is never a cached column, that quantities are integer units, and that includeZero returns net-zero pairs. This gives an agent meaningful behavioral expectations beyond the basic read-only signal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient: every sentence adds distinct value—purpose, computational provenance, filters, special case behavior, and units. It is front-loaded with the primary purpose and contains no filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential semantic details for a read-only inventory query: result granularity, filtering options, zero-balance handling, and integer units. It stops slightly short of explaining filter combination logic or the exact response shape, but for a read tool with no output schema it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates partially: it explains includeZero and names itemId, warehouseId, and locationId as filters. However, it does not mention the required workspaceId, how filters interact (e.g., AND/OR), or any value formats, leaving some parameter semantics to be inferred.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear resource—on-hand quantity by location—and adds a warehouse roll-up, distinguishing it from related inventory read tools like stock_on_hand or inventory_balance. It also specifies the data source and computational nature, which further clarifies exactly what this tool returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool: when a location-level inventory balance is needed, optionally filtered by item, warehouse, or location. It does not explicitly name alternatives or state when not to use it, but the context is strong and self-contained.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_cycle_count_statusARead-only
Der Status der Inventuren (cycle-count / stocktake status, US-J07.6): the open and recently committed J04 sessions with session_id, status, freeze_at, warehouse scope, line_count / counted_count / uncounted_count / review_required_count, committed_at, inventar_document_id, days_open and an overdue flag (open past the configured window). Optional status / warehouse_ids / overdue_only filters. Reuses the J04 inventory_stocktake_list read model. Variance qty / value are left null in Phase 1 (they need the per-session report). Useful for operational follow-up and the close-pack checklist. Reads only.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| thresholds | No | ||
| workspaceId | Yes | ||
| overdue_only | No | ||
| warehouse_ids | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is consistent with the readOnlyHint annotation and explicitly states 'Reads only.' It adds useful non-obvious context: reuse of the J04 read model, the Phase 1 limitation that variance values are null, and the overdue flag behavior. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The text is dense and information-rich but runs together as a mixed-language paragraph with slash-separated field lists. It front-loads the purpose and avoids fluff, but internal codes like 'US-J07.6' and 'J04' add jargon without much structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
It partially compensates for the missing output schema by enumerating returned fields and filters. Still, with no output schema and no parameter descriptions, it should clarify the thresholds object, valid status filter values, and the meaning of 'recently committed' to be fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description names three of the five parameters (status, warehouse_ids, overdue_only) and gives some sense of their purpose, which matters because schema coverage is 0%. However, it does not explain the required workspaceId or the nested thresholds object, leaving a meaningful gap in parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource—inventory/cycle-count status of J04 sessions—and lists the fields returned, so its purpose is concrete. It lacks an explicit action verb like 'list' or 'get' and does not explicitly distinguish itself from sibling inventory tools, though its status-focused scope is evident.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It says the tool is useful for operational follow-up and the close-pack checklist, which gives a clear intended context. It does not name alternative tools or state when not to use it; the hint that variance reporting needs the per-session report is the closest it comes to exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_ensure_default_locationA
Return the workspace default warehouse + location, creating the MAIN / DEFAULT pair if none exists yet. Idempotent and race-safe: concurrent calls produce exactly one pair. This is how a legacy D01 single-location workspace gains a default with no migration.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses idempotency and race-safety ('concurrent calls produce exactly one pair') and the side effect of creating the pair if missing. This is valuable beyond the schema and helps an agent understand the mutation and concurrency guarantees.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler. The primary action is front-loaded, followed by key behavioral traits and a usage scenario. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one parameter and no output schema, the description covers the core aspects: what it returns, when to use it, and its side effects. It doesn't specify the exact return structure, but that is minor given the clarity of 'warehouse + location'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. The single parameter workspaceId is self-evident from context ('workspace default'), but the description doesn't explicitly restate its meaning or format. Given the parameter is simple and the tool's purpose is clear, the description adds adequate context but not explicit parameter details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns the workspace default warehouse and location, and creates the MAIN/DEFAULT pair if absent. It is specific about the verb (return/ensure) and resource (warehouse/location). It doesn't explicitly name alternative tools, but the unique ensure-and-return behavior distinguishes it from nearby siblings like location_get or warehouse_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a concrete use case: 'This is how a legacy D01 single-location workspace gains a default with no migration.' This tells the agent when to invoke it. It doesn't explicitly exclude alternatives, but the scenario is specific enough to guide selection among the many inventory/location tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_get_configARead-only
Read the workspace inventory policy: allowNegativeStock (default false). When false, an issue or transfer that would drive a balance below zero is refused with insufficient_stock.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true provided by annotations, the description adds value by explaining the default value of allowNegativeStock and the consequence when it is false (refusal with insufficient_stock). No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no filler. The action and key detail are front-loaded, and the behavioral consequence is explained in a compact way.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with one self-explanatory parameter and annotations already conveying safety, the description fully enables correct invocation. It explains the policy meaning sufficiently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Despite 0% schema description coverage, the single parameter workspaceId is self-explanatory from its name and context. The description focuses on the returned policy rather than the parameter, which is acceptable for such a simple parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads the workspace inventory policy and specifies the exact field (allowNegativeStock) with its default. The verb 'Read' differentiates it from the sibling inventory_set_config.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used to check the inventory policy, but it does not explicitly state when to use it versus alternatives. It provides useful context about the setting's behavior but no direct guidance or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_lot_traceARead-only
Die Chargen- / Seriennummern-Verfolgung (lot / serial trace, US-J07.7): given a lot_code or a serial_number (optionally disambiguated by item_id), return the current position(s) (location and remaining qty for a lot, current location and status for a serial) and the full chronological J02 movement history for that lot / serial. Supports the OR 957a duty to keep orderly, verifiable inventory records. A lot / serial that does not exist in this workspace answers a soft found:false (not a hard error, and never a cross-tenant leak). Reads only, posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| item_id | No | ||
| lot_code | No | ||
| workspaceId | Yes | ||
| serial_number | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces this with 'Reads only, posts nothing.' It adds valuable behavioral context beyond the annotation: the soft found:false response for non-existent lots/serials, the explicit 'never a cross-tenant leak' guarantee, and the scope of the returned history (chronological J02 movements). It does not detail pagination or output shape, but with no output schema and readOnlyHint present, the description carries a reasonable burden and covers the key behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph that front-loads the core purpose and then adds the regulatory context, error behavior, and read-only guarantee. Every sentence earns its place, though the regulatory citation (OR 957a) and the cross-tenant leak note add length without directly affecting invocation. It is appropriately sized for a tool with 4 parameters and no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (two alternative lookup keys, optional disambiguation, required workspaceId, no output schema), the description covers the essential invocation context: what inputs are accepted, what is returned, the soft-failure mode, and the read-only safety profile. It does not specify pagination, sorting, or the exact structure of the movement history, but those are reasonable gaps for a description of this length and are partially covered by the 'full chronological J02 movement history' phrasing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does: it explains that lot_code or serial_number are the primary lookup keys, that item_id optionally disambiguates, and that workspaceId is the required scope. It does not explain the exact format of lot_code/serial_number or whether both can be provided together, but it adds meaning beyond the bare schema property names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('return'), a specific resource ('current position(s) and full chronological J02 movement history'), and the input keys (lot_code or serial_number, optionally disambiguated by item_id). It also names the regulatory purpose (OR 957a) and explicitly distinguishes itself from a hard error or cross-tenant leak. This clearly differentiates it from siblings like lot_search, serial_get, inventory_movement_history, and stock_on_hand.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when to use it: when you need current position plus full movement history for a lot or serial. It also gives an exclusion: a non-existent lot/serial returns found:false rather than a hard error. However, it does not explicitly name alternative tools for simpler lookups (e.g., lot_get, serial_get, inventory_movement_list) or state when NOT to use it in favor of those, so it misses the explicit when-not/alternatives guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_low_stockARead-only
Die Liste tiefer Bestaende (low stock, US-J07.2): every stock-tracked item whose on-hand has fallen to or below its D00 reorder point, ordered by severity (deepest shortfall, then lowest days-of-cover, first). Each row carries current_qty, reorder_point, shortfall_qty, last_movement_at, a simple average daily usage derived from recent outbound movements and an estimated days-of-cover; when there is not enough usage history the cover is null with warning usage_history_insufficient. Optional location_ids / warehouse_ids narrow the on-hand, min_shortfall filters, include_zero keeps zero-reorder items. safety_stock / suggested_reorder_qty are null until a later item-master extension adds the columns. Empty is an ok empty list. Pure derivation over the ledger; posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| thresholds | No | ||
| workspaceId | Yes | ||
| include_zero | No | ||
| location_ids | No | ||
| min_shortfall | No | ||
| warehouse_ids | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the readOnlyHint annotation: it states 'Pure derivation over the ledger; posts nothing,' clarifies that insufficient history yields null days-of-cover with a 'usage_history_insufficient' warning, notes that empty is an acceptable empty list, and flags safety_stock/suggested_reorder_qty as null until a future extension. These are concrete behavioral disclosures.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but information-dense, front-loading the core definition, then ordering, row fields, filters, and edge cases. The cryptic 'US-J07.2' and German phrase add little, but overall structure is reasonable for a tool with six parameters and no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because there is no output schema, the description compensates by listing row fields (current_qty, reorder_point, shortfall_qty, etc.), ordering, null behavior, and empty-list semantics. The unexplained 'thresholds' parameter prevents a perfect score, but the rest of the behavior is well covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate, and it does explain location_ids, warehouse_ids, min_shortfall, and include_zero. However, the 'thresholds' object parameter is entirely unexplained, and workspaceId is mentioned only as a required schema field, leaving a clear gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource and result set: 'every stock-tracked item whose on-hand has fallen to or below its D00 reorder point' with a defined ordering. However, it does not differentiate itself from the near-identically named sibling 'stock_low_stock', so an agent may struggle to choose between them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives such as inventory_reorder_candidates, stock_on_hand, or the similarly named stock_low_stock. The description only explains optional filter parameters, not the selection context that distinguishes this tool from siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_moveA
Record ONE append-only inventory movement, the only way on-hand ever changes (on-hand is always the live SUM of movements, never a stored column). movementType is one of opening|receipt|issue|transfer_out|transfer_in|adjustment|return|scrap; qty is signed integer units (opening/receipt/transfer_in must be positive, issue/transfer_out/scrap negative, adjustment/return carry the caller's sign) and must be non-zero (invalid_qty). unitCostMinor is the integer-Rappen cost snapshot for later valuation. The item must be stockable (item_not_stockable) and, when its tracking_mode demands it, carry a lotId (lot_required) and/or serialId (serial_required); a serial movement is a unit of one, so an INBOUND movement for a serial the workspace already holds is refused with serial_already_in_stock before any row is written (issue, return or scrap it first; use inventory_transfer to relocate it). When allow_negative_stock is false (default) an issue that would drive the balance below zero is refused with insufficient_stock before any row is written. Rows are immutable; a correction is another movement, never an edit.
| Name | Required | Description | Default |
|---|---|---|---|
| qty | Yes | ||
| lotId | No | ||
| itemId | Yes | ||
| serialId | No | ||
| locationId | Yes | ||
| description | No | ||
| workspaceId | Yes | ||
| movementType | Yes | ||
| effectiveDate | Yes | ||
| unitCostMinor | No | ||
| idempotencyKey | Yes | ||
| sourceDocumentId | No | ||
| sourceDocumentType | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so admirably. It discloses append-only immutability, row-level refusal conditions (invalid_qty, item_not_stockable, lot_required, serial_required, serial_already_in_stock, insufficient_stock), sign conventions per movement type, and the cost-snapshot semantics of unitCostMinor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a dense paragraph, but every sentence carries essential operational rules. It front-loads the core invariant and then methodically covers movement types, quantities, costs, tracking requirements, and error conditions. Slightly long, but no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the high parameter count (13) and the absence of annotations and output schema, the description is remarkably complete. It covers the key invariants, all error paths mentioned, and tracking-mode logic. Minor gaps include explicit idempotency semantics and return value expectations, but the tool is a write operation and the description provides enough for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains qty signing rules, movementType enumeration values, unitCostMinor purpose, and lotId/serialId conditional requirements. It does not explain idempotencyKey, effectiveDate, or sourceDocument fields, but the most complex and error-prone parameters are well covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Record ONE append-only inventory movement' and states it is 'the only way on-hand ever changes', a specific verb+resource with a strong scope statement that distinguishes it from sibling inventory tools. It clearly positions this as the core mutation point for on-hand quantity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes when to use this tool ('only way on-hand changes') and gives an explicit exclusion: for relocating a serial, use inventory_transfer instead. It does not contrast with inventory_adjust, goods_receipt, or stock_move, but the core usage boundary is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_movement_getARead-only
Read one movement by id, with its full cost snapshot, type, effective date, location, lot/serial, source document link, description and actor. A foreign or unknown id is not_found, never cross-tenant data.
| Name | Required | Description | Default |
|---|---|---|---|
| movementId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the read-only nature is covered. The description adds valuable behavioral context: a foreign or unknown id returns not_found, and data never crosses tenant boundaries. This goes beyond annotations to clarify error handling and security guarantees.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, front-loaded with the core action ('Read one movement by id') followed by a succinct list of returned fields and an error/tenant note. Every clause adds value without unnecessary filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple get-by-id tool with no output schema, the description lists the key return fields and error behavior. It does not describe the exact response envelope or mention authentication, but given the low complexity and the explicit field list, it is sufficiently complete for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameters. It only hints at 'by id' without specifying that movementId is the id or explaining the role of workspaceId. It does not mention that workspaceId is required for tenant scoping, leaving parameter semantics largely unclear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read') and resource ('one movement by id'), and enumerates the exact fields returned (cost snapshot, type, effective date, location, lot/serial, source document link, description, actor). It also clarifies error behavior (not_found) and tenant isolation, distinguishing it from list/history siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when you have a specific movement id, but does not explicitly state when to use this versus alternatives like inventory_movement_list or inventory_movement_history. No exclusions or alternative names are mentioned, but the purpose is clear enough to infer the primary use case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_movement_historyARead-only
Die Bewegungshistorie (movement history, US-J07.4): the ordered, filterable J02 movement rows for one item (required), optionally narrowed by location, lot_code, serial_number and a from_date / to_date range, with a server-computed running on-hand balance that closes to the live inventory_stock_position qty for the same filter (spec tripwire). Reuses the J02 inventory_movement_list read model, so history never becomes a second projection. Supports limit / offset pagination; an item / location / lot / serial that does not exist is not_found; from_date after to_date is invalid_date_range. Empty history is an ok empty list. Reads only, posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| item_id | Yes | ||
| to_date | No | ||
| lot_code | No | ||
| from_date | No | ||
| location_id | No | ||
| workspaceId | Yes | ||
| serial_number | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, but the description adds substantial behavioral detail: it is 'Reads only, posts nothing' (reinforcing the annotation), spells out exact error conditions ('not_found' for nonexistent item/location/lot/serial, 'invalid_date_range' when from_date after to_date), describes the empty response as 'ok empty list', and documents pagination with limit/offset. It also explains a guarantee that the running balance 'closes to the live inventory_stock_position qty' (a spec tripwire), which is valuable behavioral context an agent would need. This goes beyond the annotation without contradicting it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph but packed with information. It front-loads the purpose and then details filters, pagination, errors, and read-only nature. Each sentence earns its place; there is little fluff. The inclusion of internal codes like 'US-J07.4' and 'spec tripwire' adds a minor overhead but does not detract from the overall structure. It is longer than a simple sentence but justified by the complexity of the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (9 parameters, filters, pagination) and the absence of an output schema, the description covers a lot: filter semantics, pagination, error conditions, and the running balance guarantee. However, it fails to describe the structure of the returned data. An agent cannot know what fields the movement rows contain or how the running balance is represented in the response. Additionally, specifics like date format and pagination defaults are missing. These gaps would hinder correct invocation and interpretation of results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 9 parameters with 0% description coverage, so the description must bear the burden of explaining them. It does mention the core filters ('optionally narrowed by location, lot_code, serial_number and a from_date / to_date range') and states that item_id is required. However, it uses 'location' while the schema property is 'location_id', which could cause confusion. It does not explain workspaceId at all, nor the format of date strings, nor any defaults for limit/offset. The description gives high-level semantics but lacks precise mapping to each parameter and their formats, a significant gap given zero schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description begins with a clear statement of the tool's function: 'the ordered, filterable J02 movement rows for one item (required)' and explicitly lists the filter dimensions (location, lot_code, serial_number, date range). It also distinguishes itself by adding a 'server-computed running on-hand balance' that closes to inventory_stock_position, which is a unique capability not present in sibling tools like inventory_movement_list. The read-only nature is stated, and the phrase 'Reuses the J02 inventory_movement_list read model' positions it relative to an existing sibling without confusion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on what the tool does and when it might be used (filtered history with running balance), but it does not explicitly state when to use this tool instead of others. It mentions 'Reuses the J02 inventory_movement_list read model' but does not elaborate on the difference or when to choose one over the other. There is no mention of an alternative tool or exclusions. Thus, while the intended usage is implied, explicit guidance is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_movement_listARead-only
The chronological movement history (newest first), filterable by itemId, locationId, lotId, serialId, a movementType list, a fromDate/toDate range, sourceDocumentType/sourceDocumentId and transferGroupId, with limit/offset pagination. When a single itemId is filtered each row also carries a server-computed runningBalance (oldest to newest). History is never purged.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| lotId | No | ||
| itemId | No | ||
| offset | No | ||
| toDate | No | ||
| fromDate | No | ||
| serialId | No | ||
| locationId | No | ||
| workspaceId | Yes | ||
| movementType | No | ||
| transferGroupId | No | ||
| sourceDocumentId | No | ||
| sourceDocumentType | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description adds valuable behavioral context: newest-first ordering, a conditional runningBalance when a single itemId is filtered, and a never-purged history guarantee. These details meaningfully shape expectations without contradicting the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with no filler. The core behavior is front-loaded, filters are consolidated into one list, and the special conditional behavior and retention guarantee each earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with no output schema, the description covers ordering, filtering, pagination, a conditional computed field, and retention. It does not mention how filters combine (AND vs OR) or the response field structure, but these are minor given the strong overall coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden. It lists nearly all filterable parameters (itemId, locationId, lotId, serialId, movementType list, fromDate/toDate, sourceDocumentType/sourceDocumentId, transferGroupId, limit/offset) and explains the special runningBalance condition. It does not define value formats or enums, but it adds substantial meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource as chronological movement history and implies the 'list' verb. It is specific about ordering and the core resource, but it does not differentiate from the sibling tool inventory_movement_history, which could plausibly overlap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance about when to choose this tool over alternatives. With siblings like inventory_movement_history and inventory_stock_position present, the agent is left to infer the appropriate use case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_on_hand_by_lotARead-only
On-hand quantity broken down by lot x location: a pure read model over the OP2 stock_movement ledger (SUM of signed qty per lot x location, joined to the lot master), never a cached column. Filter by itemId, locationId, lotId, a status list and an expiryBefore cutoff; includeZero returns pairs that net to zero. Empty until J02 movements carry a lot_id.
| Name | Required | Description | Default |
|---|---|---|---|
| lotId | No | ||
| itemId | No | ||
| status | No | ||
| locationId | No | ||
| includeZero | No | ||
| workspaceId | Yes | ||
| expiryBefore | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explains the computation method (SUM of signed qty from the ledger), that it is never cached, the includeZero parameter causes zero-net pairs to be returned, and the data caveat about J02 movements. This adds significant behavioral context beyond the readOnlyHint annotation, making it transparent about how results are computed and when data may be empty.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, two sentences, with the core purpose and filtering capabilities front-loaded, and the caveat as an afterthought. No wasted words, every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool is a read-only report with seven parameters, the description covers the computation logic, filters, and a critical data caveat. There is no output schema to explain return shape, but the description's detail on what is returned (lot x location pairs, includeZero behavior) is sufficient for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explains the purpose of includeZero (returns pairs that net to zero) and lists filter dimensions (itemId, locationId, lotId, status list, expiryBefore), which adds meaning to the parameters. It does not detail formats for status or expiryBefore, but the coverage is reasonable given the schema has no descriptions, so the description compensates well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is a read-only report of on-hand quantity grouped by lot and location, derived from the OP2 stock_movement ledger. It distinguishes itself from other inventory tools like stock_on_hand and inventory_balance_by_location by specifying the lot x location granularity and the ledger source.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions it is a pure read model and notes a caveat about J02 movements to guide on data availability, but it does not explicitly state when to use this tool over alternatives like stock_on_hand or inventory_balance_by_location. The sibling list suggests these exist, so explicit guidance would improve it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_reason_archiveA
Soft-archive a reason code (isActive = 0): it disappears from pickers but stays queryable for historical joins, and new adjustments can no longer cite it. Never a destructive delete. A foreign or unknown id is not_found. Idempotent under idempotencyKey.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It thoroughly explains the soft-archive semantics (isActive=0, picker visibility, historical queryability, and inability to cite in new adjustments), explicitly states it is not destructive, describes the not_found error for foreign/unknown ids, and notes idempotency under idempotencyKey. This is comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, with each sentence serving a distinct purpose: it states the primary effect, clarifies it is not destructive, and covers error and idempotency behavior. It is front-loaded with the most important information and contains no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The operation semantics are well covered, but the description omits essential parameter explanations. Given that the schema provides no parameter descriptions and the tool has three required parameters, the description should have explained at least what id and workspaceId are. While the core behavior is clear, the missing parameter context leaves the tool incomplete for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage and the description provides no explanation for any of the three parameters. It only mentions idempotencyKey in passing without defining it, and does not clarify what workspaceId or id refer to. The description fails to compensate for the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the operation (soft-archive), the target resource (reason code), and the specific effects (disappears from pickers, stays queryable for historical joins, new adjustments can no longer cite it). It also explicitly distinguishes from a destructive delete, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides context on what the operation does but does not explicitly compare to alternatives or state when to use it over other archive tools like archive_item or inventory_reason_update. It does mention it is not a destructive delete, which gives some guidance, but no explicit routing to this tool vs. others.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_reason_createA
Create a workspace-scoped adjustment reason code. code is upper-normalised and unique per workspace, case-insensitively (duplicate_code on collision). category is one of shrinkage, damage, found, count_variance, obsolescence, theft, quality, correction, reversal, system, other. requiresNote forces a non-empty note on any adjustment citing this code; defaultForStocktake marks it selectable by a stocktake commit. Idempotent under idempotencyKey.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | ||
| name | Yes | ||
| category | Yes | ||
| description | No | ||
| workspaceId | Yes | ||
| requiresNote | No | ||
| idempotencyKey | Yes | ||
| defaultForStocktake | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral disclosure burden and does so thoroughly. It reveals 'code is upper-normalised and unique per workspace, case-insensitively (duplicate_code on collision)', enumerates allowed categories, explains effects of requiresNote and defaultForStocktake, and states 'Idempotent under idempotencyKey.' This is far beyond a generic create description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four dense sentences, front-loaded with the core action ('Create a workspace-scoped adjustment reason code') and then ordered constraints: code semantics, category values, boolean flags, and idempotency. Every sentence adds necessary information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even in the absence of annotations and an output schema, the description provides everything an agent needs to invoke the tool correctly: scope, naming behavior, uniqueness rules, allowed category values, flag effects, and idempotency guarantees. No call-blocking information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does for the most important parameters: code normalization/uniqueness/collision behavior, category allowed values, requiresNote/defaultForStocktake semantics, and idempotencyKey. workspaceId, name, and description are not explicitly explained, though their meaning is largely inferable from the schema and the phrase 'workspace-scoped.'
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Create a workspace-scoped adjustment reason code.' It clearly distinguishes this tool from the sibling reason tools (update, archive, list, get) by emphasizing creation and workspace scoping. The normalization and uniqueness details reinforce the create-specific purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: when a new workspace-scoped adjustment reason code is needed. It does not explicitly name alternatives or exclusion conditions, but the create/update/archive sibling split makes the intended selection unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_reason_getARead-only
Read one reason code by id. A foreign or unknown id is not_found.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds valuable behavioral context by specifying that a foreign or unknown id results in not_found, which helps the agent anticipate error behavior without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences deliver the essential meaning and error behavior with no redundant wording. The core action is front-loaded, and the not_found clarification earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-by-id tool with readOnlyHint and no output schema, the description is nearly complete: it names the action, the target resource, the key parameter, and a notable failure mode. It only lacks a minimal note on workspaceId semantics, but the overall guidance is sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of explaining parameters. It meaningfully clarifies that 'id' selects a single reason code and that unrecognized ids yield not_found. However, it does not explain the role of workspaceId, leaving one of the two required parameters semantically unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Read one reason code by id.' This clearly identifies the tool as a single-entity fetch, distinguishing it from sibling tools like inventory_reason_list, inventory_reason_create, and inventory_reason_update.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'by id' phrasing implies this tool is for retrieving a single reason code rather than listing multiple ones, but it does not explicitly state when to use this tool versus alternatives or mention any exclusions. The usage context is implied, not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_reason_listARead-only
List reason codes ordered by code, filterable by activeOnly and category. Returns each code with its category, requiresNote / defaultForStocktake flags and active lifecycle.
| Name | Required | Description | Default |
|---|---|---|---|
| category | No | ||
| activeOnly | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, and the description adds useful behavioral detail beyond that: results are sorted by code, filterable by activeOnly and category, and each returned code includes category, flags, and lifecycle state. No contradiction with the read-only annotation is present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence that front-loads the core action and packs in ordering, filtering, and return content. The phrase 'active lifecycle' is slightly vague, but the overall structure is tight and informative without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with simple parameters and no output schema, the description covers the essential context: what is listed, how it is ordered, what filters exist, and what each result contains. It lacks only minor details such as pagination or explicit default behavior for activeOnly, but nothing critical is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does clarify the meaning of 'activeOnly' and 'category' as filters and 'ordered by code' as ordering context, but it does not mention the required 'workspaceId' or provide any format/enum detail. The compensation is partial but adequate for these self-explanatory parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') with a clear resource ('reason codes') and adds distinctive detail: ordering by code, available filters, and returned fields. This clearly differentiates it from sibling tools like inventory_reason_get, inventory_reason_create, and inventory_reason_update.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the tool's purpose clear but does not explicitly state when to prefer it over alternatives or mention any exclusions. Usage is implied by the tool's name and the list/filter framing, but no direct sibling comparison or when-not-to-use guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_reason_updateA
Update a reason code. The code itself is immutable (it is the classifier historical rows join on); name, description, requiresNote, defaultForStocktake and isActive may change. Setting isActive false is the same soft-archive inventory_reason_archive performs. A foreign or unknown id is not_found. Idempotent under idempotencyKey.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| name | No | ||
| isActive | No | ||
| description | No | ||
| workspaceId | Yes | ||
| requiresNote | No | ||
| idempotencyKey | Yes | ||
| defaultForStocktake | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full behavioral burden and does so strongly: it discloses the immutable code constraint, the exact mutable fields, the soft-archive equivalence of isActive=false, the not_found error for unknown ids, and idempotency under idempotencyKey. This is precisely the kind of behavioral context an agent needs before invoking a mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences cover scope, mutation constraints, archive equivalence, error behavior, and idempotency in under 250 characters with no filler. The most important constraint (immutable code is front-loaded, and every sentence contributes new information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter update tool with no annotations and no output schema, the description covers the critical operational facts: which field is the immutable key, which fields may change, what isActive=false means, error behavior for bad ids, and idempotency. Nothing essential for calling this tool correctly appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does meaningfully: it explains that code is immutable, identifies which fields (name, description, requiresNote, defaultForStocktake, isActive) are updatable, and defines idempotencyKey behavior. It does not explain the semantics of requiresNote or defaultForStocktake in detail, but it adds substantial meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Update a reason code') and immediately distinguishes itself from the archive sibling by declaring the code immutable and enumerating exactly which fields are mutable. This allows an agent to separate it from inventory_reason_archive and inventory_reason_create without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context by naming the update-versus-archive relationship: setting isActive false is the same soft-archive that inventory_reason_archive performs. It does not explicitly state 'use archive instead when...' but the alternative and its behavioral equivalence are clearly implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_reconciliation_checkARead-only
The hard reconciliation check a period close calls before it hard-locks a period (YYYY-MM). Returns { status: balanced } when a valuation has been posted at the period end and every inventory control account shows delta 0. Returns a structured valuation_missing error when a non-zero inventory value exists at the period end but no run was ever posted, or a reconciliation_drift error listing the offending accounts and amounts when a delta remains. The period cannot be hard-locked until the difference is explained or a corrective valuation is posted and re-checked. Writes nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
While the annotation already declares readOnlyHint true, the description reinforces 'Writes nothing' and additionally reveals the behavioral gate: the period cannot be hard-locked until the difference is resolved. It also explains all possible return structures, which is significant beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and well-structured: it opens with the tool's purpose, then enumerates return outcomes, then the locking consequence, and ends with a clear side-effect statement. Every sentence contributes value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without an output schema, the description fully explains return values, error conditions, and preconditions. It covers everything an agent needs to know to call this tool appropriately and interpret the result, especially given its simple two-parameter interface.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and there are only two parameters. The description adds critical meaning for 'period' by specifying the format (YYYY-MM), which is essential for correct invocation. WorkspaceId is not described, but it is a common standard parameter. This partially compensates for missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it is the hard reconciliation check called by period close before hard-locking a period. It clearly distinguishes itself from other inventory tools by describing the exact return states (balanced, valuation_missing, reconciliation_drift) and its role in the locking process.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool (during period close to verify a period can be hard-locked), but it does not explicitly mention alternatives or exclusion criteria. It names the context but not other similar tools like inventory_reconciliation_report, so full discrimination is not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_reconciliation_reportARead-only
The OP11 reconciliation: per inventory control account, the live sub-ledger valuation at the cut-off (period YYYY-MM or asOf date), the GL balance of the same account at the same cut-off, the delta between them, and a status of balanced (delta 0), drift (a delta remains though a run was posted at the cut-off, e.g. an external GL-only posting) or unposted (the live valuation differs from the GL and no run has closed the gap; the delta is what a new run would post). Returns balanced/drift/unposted counts and the total unposted delta. Optional accountIds narrows the accounts. Writes nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| period | No | ||
| accountIds | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces this with 'Writes nothing.' It adds valuable behavioral detail beyond the annotation by explaining what drift vs unposted means, what counts are returned, and that the delta represents what a new run would post. This gives the agent a solid mental model of report semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause earns its place, defining the report scope, status semantics, return summary, optional filter, and side-effect safety. It is front-loaded with the report identity and delivers needed definitions without padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the full burden of explaining return values; it does so clearly, including counts and total unposted delta. It also covers cutoff selection, status meanings, optional filtering, and the read-only nature, making it complete enough for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description compensates well for most parameters: it explains period format ('YYYY-MM'), the asOf alternative, and that accountIds narrows the accounts. The required workspaceId is left to the schema and not described, which is a minor gap, but the most ambiguous parameters are clarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a very specific resource: OP11 inventory reconciliation per inventory control account, comparing sub-ledger valuation to GL balance with delta and status. It clearly differentiates from sibling reporting tools by defining the exact output fields and statuses (balanced/drift/unposted).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes it clear when this report is relevant (reconciling inventory sub-ledger to GL at a cutoff) and how to narrow scope with accountIds, but it does not explicitly name alternative tools or state when not to use this one. Usage is implied rather than explicitly contrasted with siblings like inventory_reconciliation_check or inventory_valuation_report.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_reorder_candidatesARead-only
Die Nachbestell-Kandidaten (reorder candidates, US-J07.10): the subset of the low-stock list with a POSITIVE shortfall, each with a suggested_qty (the shortfall in Phase 1), current_qty, reorder_point, estimated days-of-cover and a ready-to-use payload an agent can hand to an I00 requisition or a purchase-order tool. preferred_supplier_id is null until a supplier-preference extension lands. Advisory ONLY: no document is created and nothing is posted. Optional warehouse_ids / location_ids scope. Empty is an ok empty list.
| Name | Required | Description | Default |
|---|---|---|---|
| thresholds | No | ||
| workspaceId | Yes | ||
| location_ids | No | ||
| warehouse_ids | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds significant behavioral detail beyond the readOnlyHint annotation: it states 'Advisory ONLY: no document is created and nothing is posted,' discloses that preferred_supplier_id is null until an extension, and notes that an empty list is acceptable. This provides valuable context about side effects, null ability, and return behavior that the annotation alone does not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense yet efficient, packed with multiple key facts in a single paragraph. It is front-loaded with the core purpose and flow, but includes an internal reference (US-J07.10) that may be noise to an agent. Overall, each sentence earns its place, though a bit of trimming could improve clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the output fields, the advisory nature, the optional scope, and the empty-list behavior, which is substantial for a read-only tool without an output schema. It does not explain the thresholds parameter, or any pagination/limits, but these are minor given the tool's simplicity. The description is largely complete for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explains that warehouse_ids and location_ids are optional scope filters, which is useful given the schema has no parameter descriptions. However, it does not mention the required workspaceId or the thresholds parameter at all, and the schema allows additionalProperties, so the description leaves significant parameter ambiguity. With schema description coverage at 0%, the description partially compensates but is not comprehensive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns the subset of the low-stock list with a positive shortfall, names the specific output fields (suggested_qty, current_qty, reorder_point, etc.), and explicitly marks it as advisory-only. It distinguishes itself from the full low-stock list by being the filtered subset, and specifies it provides a payload for downstream I00/PO tools. This is a precise verb+resource+filter definition.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use this tool: when you need reorder candidates with a positive shortfall, and it warns it is advisory only (no creation or posting), indicating not to use it for transactional actions. It mentions it hands off to requisition or PO tools, implying a follow-up workflow. However, it does not explicitly name sibling alternatives (e.g., inventory_low_stock) or state 'use this instead of X', so some inference about the full list tool is needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_set_configA
Set the workspace negative-stock posture. allowNegativeStock true disables the insufficient_stock guard (a warning posture for backorder workflows); false (default) enforces it. Plain policy: posts no journal entry and never rewrites history, so flipping the flag changes only future writes.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| allowNegativeStock | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It explicitly discloses that the operation posts no journal entry and never rewrites history, and that it only affects future writes. This is a meaningful behavioral guarantee beyond typical expectations, though it does not cover idempotency or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the purpose, then the effect and a critical safety guarantee. Every sentence earns its place with no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple configuration setter, the description covers the main parameter, the behavioral effect, and the non-destructive nature. It does not mention return values or error conditions, but for a setter this is acceptable. It could note that the operation is idempotent, but overall adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It thoroughly explains allowNegativeStock (true disables guard, false enforces), but does not explain workspaceId (obvious) or idempotencyKey (common but not described). Partial compensation; could be improved by noting idempotencyKey's role in deduplication.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Set') and resource ('workspace negative-stock posture'), and clearly distinguishes from sibling inventory_get_config by focusing on setting the configuration. The description explains the exact effect of the allowNegativeStock flag, leaving no ambiguity about the tool's role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context on when to use this tool: to flip the negative-stock posture for backorder workflows. It implies the alternative (inventory_get_config) for reading the current posture, though it does not explicitly name alternatives or exclusions. The behavioral note that it only affects future writes guides safe usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_slow_moversARead-only
Die Ladenhueter-Ansicht (slow movers / aging, US-J07.8): stock-tracked items with a positive on-hand and no OUTBOUND movement inside a window (min_days_no_movement, default 90), ordered by extended value descending, each with current_qty, last_movement_at, last_outbound_at, days_idle, unit cost and extended_value_rappen (from the pure J03 book value). Optional min_value_rappen and location_ids narrow the list. This surfaces the high-value stagnant stock that lower-of-cost-or-market thinking (OR 960c) reviews; it performs NO write-down itself. Empty is an ok empty list. Reads only.
| Name | Required | Description | Default |
|---|---|---|---|
| thresholds | No | ||
| workspaceId | Yes | ||
| location_ids | No | ||
| min_value_rappen | No | ||
| min_days_no_movement | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces this with 'Reads only' and 'performs NO write-down itself.' It adds useful behavior beyond the annotation: empty results are an acceptable empty list, the valuation basis is the pure J03 book value, and results are sorted by extended value descending. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient: it front-loads the view name, then covers criteria, ordering, returned fields, filters, and empty-result behavior in a compact block. Some domain-specific codes like US-J07.8 and OR 960c may add noise, but every sentence contributes useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by listing each returned field, the ordering, the default window, and the empty-list expectation. The main gap is the unexplained 'thresholds' object, which makes the parameter surface slightly incomplete. Overall it is strong but not maximal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the semantics. It defines min_days_no_movement with a default of 90, explains that min_value_rappen and location_ids narrow the list, and enumerates the computed output fields. The 'thresholds' object and required workspaceId remain opaque, so it does not fully cover all parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise function: list stock-tracked items with positive on-hand and no outbound movement in a window, ordered by extended value. It distinguishes itself from inventory siblings by naming the specific 'slow movers / aging' view and by clarifying it performs no write-down. An agent can tell what this tool computes without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives the intended business use: surfacing high-value stagnant stock for lower-of-cost-or-market review. It also explicitly notes that the tool 'performs NO write-down itself' and is read-only, which prevents misuse. It does not name alternatives such as inventory_low_stock, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_stock_positionARead-only
Die Bestandsposition (stock position, US-J07.1): what is on hand, where, and (optionally) what it is worth, in one structured call. Sum the append-only J02 movement ledger into (item, location) positions under an optional filter (item_ids, location_ids, warehouse_ids, lot_codes, serial_numbers, only_positive), optionally rolled up by group_by (none | item | location | warehouse). On-hand is always SUM(stock_movement.qty), never a stored column, so a position can never disagree with the ledger it sums. With include_valuation the extended value comes from the pure J03 valuation preview, integer Rappen, so it equals a direct valuation call for the same (item, location). Returns rows plus a totals footer (count, qty, value_rappen) and stable limit / offset pagination; an empty match answers ok with message no_positions_matching. A foreign item / location / warehouse id is not_found (tenant isolation). Posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| filter | No | ||
| offset | No | ||
| group_by | No | ||
| workspaceId | Yes | ||
| include_valuation | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the readOnlyHint annotation by disclosing that positions are computed from the append-only ledger rather than stored, guaranteeing consistency. It also details the valuation source (pure J03 preview), stable pagination, empty-result behavior (ok with no_positions_matching), and foreign-id error semantics (not_found with tenant isolation).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence adds a distinct fact: purpose, computation method, filters, grouping, valuation origin, pagination, empty/error cases, and read-only nature. The description is front-loaded with the core purpose and contains no filler, making it dense yet efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given six parameters, no output schema, and only readOnlyHint, the description covers the major behavioral contracts: data derivation, filtering, grouping, valuation, pagination, and error semantics. It does not specify the exact output row structure, but it mentions rows plus a totals footer, so it is nearly complete for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by explaining the filter subfields (item_ids, location_ids, warehouse_ids, lot_codes, serial_numbers, only_positive), group_by enum values, include_valuation semantics, and stable limit/offset pagination. It does not detail workspaceId or provide defaults, but the key parameters are meaningfully documented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states precisely that it sums the append-only J02 movement ledger into (item, location) positions, with optional filters and grouping, and optionally computes valuation. This clearly defines the tool's resource and operation, distinguishing it from siblings like stock_on_hand or inventory_balance by combining position and valuation in one structured call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when you need on-hand quantities by location, optionally filtered/grouped, and optionally valued. It gives clear context for its intended use, though it does not explicitly name alternative tools or exclusion conditions, so it falls short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_stocktake_approve_linesA
Approve review-required (or counted) lines so they may be committed. lineIds is an explicit array of line ids or the string 'all_review_required'. An uncounted line cannot be approved (uncounted_or_unapproved_lines). A non-open session is stocktake_not_open. Idempotent under idempotencyKey.
| Name | Required | Description | Default |
|---|---|---|---|
| lineIds | Yes | ||
| sessionId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses idempotency under idempotencyKey, error conditions for uncounted lines and non-open sessions, and explains the special sentinel value 'all_review_required'. It does not mention permissions or side effects beyond enabling commit, but it covers key behavioral aspects that the schema does not.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loaded with the primary purpose, followed by key parameter detail and then error conditions. Every sentence contributes new information; there is no redundancy or fluff. It is appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers preconditions, special parameter values, and idempotency behavior. It does not describe the return value or output, but no output schema is present, so the description could arguably mention what happens on success (e.g., lines become committed). It also does not discuss permissions, but for a stocktake approval action, the covered aspects are sufficient for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the meaning of lineIds, including the special string value 'all_review_required', and clarifies idempotencyKey via the idempotency statement. workspaceId and sessionId are left to convention but are self-explanatory in context. The description adds meaningful semantics for the most complex parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb-resource pair: 'Approve review-required (or counted) lines' and states the purpose 'so they may be committed'. It names the specific action and resource, and the context distinguishes it from related stocktake operations like counting or committing. It does not rely on the tool name alone, so purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides usage context: it is used for lines that are review-required or counted, and it states preconditions (uncounted lines cannot be approved, session must be open). However, it does not explicitly contrast with sibling tools like inventory_stocktake_commit or inventory_stocktake_count, nor does it say when NOT to use this tool in favor of another. It implies the context but lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_stocktake_cancelA
Cancel an open or review session. No movements are written; the frozen snapshot is retained for audit but never re-used. A committed session cannot be cancelled (already_committed). Idempotent under idempotencyKey.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| sessionId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses key behaviors: no movements are written, the frozen snapshot is retained for audit but never re-used, and the operation is idempotent. This goes well beyond a bare 'cancel' statement, though it does not specify what happens to the session status after cancellation or detail error codes beyond the mentioned already_committed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no redundancy. The main action is front-loaded ('Cancel an open or review session'), followed by concise behavioral constraints and the idempotency note. Every clause adds necessary information, making it efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a cancel operation with no output schema or annotations, the description covers the essential semantics: valid target states, side-effect guarantees (no movements), audit retention, the committed-state error, and idempotency. It does not describe the return value or post-cancellation session status, but these are secondary for a cancel tool and the description is otherwise sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only elaborates on idempotencyKey (idempotent under it), leaving workspaceId, sessionId, and reason unexplained. The parameter names are self-explanatory in context, and the idempotency note adds value, but the description does not fully fill the gap left by the empty schema descriptions for three of four parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Cancel') and resource ('open or review session'), clearly distinguishing it from sibling stocktake operations like inventory_stocktake_count, inventory_stocktake_commit, and inventory_stocktake_open. The phrase 'open or review' further delimits the target state, leaving no ambiguity about what it operates on.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly excludes committed sessions ('A committed session cannot be cancelled (already_committed)'), giving a clear when-not-to-use condition. It also mentions idempotency under idempotencyKey, which is a usage-relevant property. However, it does not name alternative tools (e.g., a revert/reverse sibling) or outline broader selection criteria beyond the committed-state exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_stocktake_commitA
Commit a reviewed session: for every line with a non-zero variance (counted - book) mint ONE J02 inventory_move (movement_type adjustment, signed qty, dated at freezeAt, ref_kind stocktake_session) and record it on the line, then mark the session committed, all atomically. No quantity is written directly; on-hand is the SUM over the minted movements. Blocked while any line is pending or review_required (uncounted_or_unapproved_lines). Refused if the freeze date falls in a locked or sealed period (period_locked). Idempotent under idempotencyKey (a replay returns the original movement ids, posting nothing); a committed session re-committed under a NEW key is already_committed. A committed session is corrected by a compensating inventory_move, never an edit.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does an excellent job. It discloses atomicity, the exact movement type and reference kind, the rule that quantities are not written directly but derived from summed movements, idempotency semantics under idempotencyKey, the already_committed error on new keys, blocking conditions, and the correction policy via compensating moves. This is far more transparent than typical tool descriptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph, but every sentence contains essential operational detail. It is front-loaded with the core action ('Commit a reviewed session') and then systematically covers the movement creation, quoting rules, blocking/refusal conditions, idempotency, and correction policy. No filler or repetition. The structure is effective for AI consumption.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of annotations and an output schema, the description provides a thorough account of behavior, preconditions, and idempotency. It does not explicitly describe the success return value (though it hints at movement ids), and it assumes familiarity with terms like 'reviewed session' and 'freezeAt'. However, it covers the key decision-making context an agent needs to invoke the tool correctly and understand its effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides no descriptions (0% coverage), so the description must compensate. It explains the idempotencyKey parameter in detail, including replay behavior and new-key behavior, and clarifies that sessionId refers to the stocktake session being committed. The workspaceId parameter is not explained, but it is a common context parameter that may not need extensive explanation. Overall, the description adds significant meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: commit a reviewed stocktake session by minting J02 inventory moves and marking the session committed. It names the specific verb ('Commit'), the resource ('reviewed session'), and the mechanism (minting inventory_move records). This distinguishes it from sibling tools like inventory_stocktake_count or inventory_stocktake_approve_lines, which handle different phases.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit preconditions and constraints: it is blocked while lines are pending or review_required, and refused if the freeze date is in a locked or sealed period. It also clarifies idempotency behavior with replayed and new keys Nag, which guides when and how to call it. However, it does not explicitly name alternative tools for other stages, so it doesn't fully spell out 'when not to use' versus specific siblings, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_stocktake_countA
Record counted quantities for one or more lines of an open session. Each line names itemId + locationId (+ optional lotId / serialId) and countedQty (integer thousandths, may be 0 for an explicit empty bin). Last write wins while the session is open. A line with zero variance auto-approves; a material variance (over threshold) becomes review_required; else counted. A blind session still hides book_qty in the response until review. A line not in the session is not_found; a non-open session is stocktake_not_open.
| Name | Required | Description | Default |
|---|---|---|---|
| lines | Yes | ||
| sessionId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does so thoroughly. It discloses the idempotency behavior ('Last write wins'), variance-based state transitions (auto-approve, review_required, counted), the blind-session hiding of book_qty, and error conditions (not_found, stocktake_not_open). This is highly transparent and goes well beyond a simple mutating action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph that front-loads the core purpose and then details behaviors. It avoids redundancy and packs substantial information efficiently. The structure is clear, though it could benefit from slight segmentation to improve scanability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description should clarify the return structure. It only hints that a response includes book_qty in blind sessions, but does not describe the full response format, per-line results, or idempotency semantics. Error cases are covered, but an agent would still lack clarity on what the tool returns on success or partial success.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains the line object parameters (itemId, locationId, optional lotId/serialId, countedQty as integer thousandths, may be 0 for empty bin), which adds value. However, it does not explain the top-level parameters (workspaceId, sessionId, idempotencyKey) beyond obvious context, leaving idempotencyKey's purpose implicit. Partial compensation for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Record counted quantities') and the specific resource (lines of an open stocktake session). It also details the line structure and expected outcomes, which makes the tool's purpose unambiguous. However, it does not explicitly differentiate itself from the sibling tool stock_stocktake_count, which appears to serve a similar role, so it misses the top mark.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides several behavioral conditions (e.g., last-write-wins, variance outcomes, blind session behavior) that imply when the tool is appropriate, but it does not explicitly state when to use this tool over alternatives like inventory_stocktake_approve_lines or stock_stocktake_count. There is no mention of prerequisites or when not to use it, leaving selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_stocktake_createA
Open a cycle-count / stocktake session that FREEZES a book-quantity snapshot from the J02 movement ledger. type is 'full' (an OR 958c Abs. 2 Inventur frozen at a balance-sheet date) or 'cycle' (an ongoing scope count without a full freeze; default 'full'). freezeAt is the ISO date the book quantities are snapshotted as-of (default today); book_qty per line is SUM(stock_movement.qty) WHERE moved_at <= freezeAt, bit-identical to inventory_balance asOf. Scope with any of warehouseId, locationIds[], itemIds[] (a foreign id is not_found, §H-TENANT); includeZeroQty forces net-zero groups into the count. blindCount hides book_qty until review. varianceQtyThreshold / variancePctThreshold (integer, default 0) classify a counted line as material: a non-zero variance exceeding either enters review_required and blocks commit until approved. Idempotent under idempotencyKey.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | ||
| notes | No | ||
| itemIds | No | ||
| freezeAt | No | ||
| abcClasses | No | ||
| blindCount | No | ||
| locationIds | No | ||
| warehouseId | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| includeZeroQty | No | ||
| variancePctThreshold | No | ||
| varianceQtyThreshold | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: it discloses the freeze behavior, exact book_qty computation, default values for type and freezeAt, idempotency under idempotencyKey, blindCount behavior, and the review_required/commit-blocking consequence of variance thresholds. This is far more than a typical tool description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence contributes; the main purpose is front-loaded. It loses a point for readability—heavy compression, legal reference, and long semicolon-separated clauses make it harder to parse than necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter tool with no output schema and no annotations, the description is unusually thorough on parameters and behavior. The main gap is that it does not describe the response shape or what the agent should expect after successful creation, which would be valuable without an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, yet the description adds meaning to nearly every parameter: type, freezeAt, warehouseId/locationIds/itemIds, includeZeroQty, blindCount, varianceQtyThreshold, variancePctThreshold, and idempotencyKey. It converts bare type names into operational semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action—'Open a cycle-count / stocktake session'—and identifies the resource and core behavior: freezing a book-quantity snapshot from the J02 movement ledger. It also distinguishes 'full' from 'cycle' counts, which differentiates this creation tool from sibling count/commit/report tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context on when to use 'full' vs 'cycle', how to scope with warehouse/location/item IDs, and how thresholds trigger reviews. It does not explicitly name sibling alternatives like inventory_stocktake_count or inventory_stocktake_commit, nor provide when-not-to-use exclusions, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_stocktake_getARead-only
Read one session with its header (type, status, freeze date, progress, variance summary, Inventar link) and its lines. A blind session hides book_qty until review. A legacy D01 stocktake id resolves read-only (legacy:true). A foreign or unknown id is not_found.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses meaningful behavioral traits: blind sessions hide book_qty until review, legacy D01 IDs resolve as read-only with legacy:true, and foreign/unknown IDs return not_found. These are specific edge cases that affect the agent's interpretation of results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is four sentences, front-loaded with the core purpose, followed by conditional behaviors. Each sentence adds distinct information without redundancy. It is appropriately concise for a read operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, the description provides a solid outline of the returned data (header fields and lines), plus critical behavioral variants (blind session, legacy, not_found). It could be more explicit about the response structure, but covers the essential context an agent needs to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of explaining parameters. It does not explicitly define workspaceId or sessionId, though 'legacy D01 stocktake id' hints that sessionId is the identifier. There is no mention of what workspaceId is for or any format expectations.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read') and resource ('one session'), and enumerates the exact header fields and lines returned. It clearly distinguishes this from sibling stocktake tools (list, count, report) by focusing on retrieving a single session with its header and lines.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is the read operation for a stocktake session, but it does not explicitly mention alternatives or when to choose this over sibling tools like inventory_stocktake_list or inventory_stocktake_report. The behavior notes (legacy resolution, blind session) give context, but no direct 'use this when...' guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_stocktake_listARead-only
List sessions (newest freeze date first), filterable by type, status (one or an array), a freeze-date from/to range and warehouseId. Set includeLegacy to also surface legacy D01 stocktakes read-only. Each row carries progress and the counted / review / total line counts.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| type | No | ||
| status | No | ||
| warehouseId | No | ||
| workspaceId | Yes | ||
| includeLegacy | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description adds value by specifying the ordering (newest freeze date first), the filtering semantics, the includeLegacy behavior (surfacing legacy D01 read-only), and the row contents (progress, counted/review/total line counts). This goes beyond the annotation without contradicting it, providing useful behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences front-load the core purpose (list sessions, newest first) and then enumerate filters and special behavior. No redundant wording; every clause adds information. The structure is efficient and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with 7 parameters and no output schema, the description covers the key aspects: purpose, ordering, filters, includeLegacy behavior, and row content. It omits pagination/limits and valid enum values, but those are not critical for initial invocation. The annotation covers read-only safety, making the description sufficient for an agent to decide when and how to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains type, status (one or an array), freeze-date from/to range, warehouseId, and includeLegacy with its purpose. It does not detail valid type values or date formats, and workspaceId is only implied as a required workspace context, but the main filters are adequately described.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists stocktake sessions, specifies the ordering (newest freeze date first), enumerates the filterable dimensions (type, status, freeze-date range, warehouseId), and notes the includeLegacy option. It distinguishes this from other stocktake tools like create/count/commit by focusing on listing, and names the legacy read-only behavior. This is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (when you need to list stocktake sessions) but does not explicitly compare to alternatives like inventory_stocktake_get (single fetch) or other list tools. It gives no 'when not to use' guidance or exclusions. While the filtering and includeLegacy hint at scenarios, it lacks explicit routing against sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_stocktake_reportARead-only
The variance report for a session (pure read, writes nothing). Returns per-line book_qty, counted_qty, variance_qty and variance_pct plus aggregate over / under / absolute variance and the count of lines exceeding threshold. book_qty is hidden on a blind session until it reaches review. A foreign or unknown session id is not_found.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotation readOnlyHint=true already indicates a read-only operation. The description adds 'writes nothing', reinforcing this, and discloses two behavioral quirks: book_qty is hidden on blind sessions until review, and an unknown or foreign session id results in 'not_found'. These go beyond the annotation and provide useful handling context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, with the primary purpose front-loaded ('variance report'), followed by return data and key behaviors. Every sentence adds value; there is no filler or unnecessary detail. It is compact and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only report with two parameters and no output schema, the description covers all essential aspects: what it returns (line items and aggregates), the blind session behavior, and error handling. The only minor omission is the definition of 'threshold', but that likely refers to a known threshold parameter elsewhere in the system. Overall, it is complete for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the description does not explain the parameters. sessionId is mentioned in context ('foreign or unknown session id') but not defined as a parameter. workspaceId is not addressed at all. The parameters are common, but the description fails to clarify their role or any constraints, leaving the agent to infer from names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a variance report for a stocktake session, specifying the exact data it returns (per-line quantities and variance metrics). It distinguishes itself from sibling tools like inventory_stocktake_count or stock_stocktake_commit by stating it is a pure read operation that produces a report. The verb 'report' and resource 'session' are specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when you need variance analysis for a stocktake session. It does not explicitly mention alternatives or exclusions, but the presence of sibling names like inventory_stocktake_get or stock_stocktake_report suggests a clear domain. The description does not say 'use this instead of X', but the context makes it obvious it is the reporting variant.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_stocktake_request_recountA
Send lines back to pending, clearing their counted quantity so they must be recounted. lineIds is the array of line ids. A non-open session is stocktake_not_open. Idempotent under idempotencyKey.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| lineIds | Yes | ||
| sessionId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and already discloses idempotency under idempotencyKey garland the error condition for a non-open session. It also explicitly states the side effect (clearing counted quantity). Missing details like reversibility or permission requirements but the key behavioral traits are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no fluff. The primary action is front-loaded, followed by the key parameter clarification, then behavioral notes. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 5 parameters, no output schema, and no annotations, the description lacks essential details about the required workspaceId and sessionId parameters, does not explain the reason parameter, and does not describe return values or response behavior. It is incomplete for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains lineIds ('array of line ids') and mentions idempotencyKey indirectly, but sessionId, workspaceId, and reason are not described at all. This is insufficient for an agent to correctly populate all parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Send lines back to pending'), the specific effect ('clearing their counted quantity'), and the goal ('so they must be recounted'). It distinguishes this tool from other stocktake siblings like inventory_stocktake_count or inventory_stocktake_commit by explicitly focusing on the recount flow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage via 'so they must be recounted' but does not explicitly state when to use this tool over alternatives (e.g., inventory_stocktake_count, inventory_stocktake_commit) or when not to use it. No exclusions or sibling comparisons are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_transferA
Move stock between two locations as an ATOMIC pair: a transfer_out at fromLocationId and a transfer_in at toLocationId, both under one transferGroupId, equal absolute quantity, summing to zero. qty is always POSITIVE (the helper applies the sign). Either both legs land or neither does. When allow_negative_stock is false a transfer that would overdraw the source is refused with insufficient_stock and nothing is written. Same tracking rules as inventory_move; a serial transfer moves one unit and re-parks it at the destination. Idempotent under idempotencyKey (a replay returns the original pair).
| Name | Required | Description | Default |
|---|---|---|---|
| qty | Yes | ||
| lotId | No | ||
| itemId | Yes | ||
| serialId | No | ||
| description | No | ||
| workspaceId | Yes | ||
| toLocationId | Yes | ||
| effectiveDate | Yes | ||
| unitCostMinor | No | ||
| fromLocationId | Yes | ||
| idempotencyKey | Yes | ||
| sourceDocumentId | No | ||
| sourceDocumentType | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden and does so excellently. It discloses atomicity, the positive-qty sign convention, failure mode when allow_negative_stock is false, idempotent replay behavior, and serial-transfer re-parking behavior. These are non-obvious traits beyond what the tool name or schema could convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place. It front-loads the core atomic-transfer concept, then covers sign convention, failure behavior, serial handling, and idempotency in compact logical order. There is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex inventory operation with no annotations and no output schema, the description supplies most of the critical operational context: atomicity, failure semantics, idempotency, and serial behavior. It still omits some parameter meanings and does not describe the response shape or permission prerequisites, but the provided context is substantially richer than typical for this tool set.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds crucial meaning for qty (always positive) and for the from/to pair, and mentions transferGroupId even though it is not in the schema. However, many parameters remain unexplained: unitCostMinor, sourceDocumentId, sourceDocumentType, description, effectiveDate format, and the lotId/serialId selection rules. The contribution is meaningful but incomplete for a 13-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource, then defines the transfer precisely as an atomic pair of transfer_out and transfer_in operations. It clearly distinguishes itself from simpler movement tools like inventory_move by emphasizing atomicity and shared transferGroupId. There is no ambiguity about what this tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the usage context clear: use when a stock transfer between two locations must be atomic and either fully succeeds or fully fails. It provides relevant conditions like allow_negative_stock and idempotency. It does not explicitly name when NOT to use it or compare it against inventory_move as an alternative, but the atomic-pair framing is strong enough context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_valuation_createA
Create a DRAFT inventory valuation run for a cut-off (asOf ISO date, or period YYYY-MM for the period end) and return it with its proposed adjusting journal for review. Computes the FULL inventory valuation through the J03 engine (each item under the method in force on that date), writes one immutable line per (item, location) that carries a quantity or a value, and stores the run total. Posts NO journal: that is inventory_valuation_post. netRealisableValues maps itemId to the OR 960c per-unit Veraeusserungswert less costs to come; where it is below cost the J03 clamp writes the line down and marks it. inventoryAccountId / changeAccountId override the default 1200 / 4200 control accounts. A run always values the complete position (a filtered run cannot hold the sub-ledger = GL identity; use inventory_valuation_report for a filtered view). A cut-off in a soft- or hard-closed period is refused with period_locked before anything is written. Idempotent: a replay of the key returns the same draft and recomputes nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| notes | No | ||
| method | No | ||
| period | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| changeAccountId | No | ||
| inventoryAccountId | No | ||
| netRealisableValues | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and it delivers extensively. It discloses the draft-only nature, J03 engine computation, immutable per-item lines, run-total storage, no journal posting, override behavior, complete-position requirement, closed-period refusal, and idempotency. This is far beyond a basic 'creates a draft' and gives an agent a thorough mental model of side effects and constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place. It fronts the core purpose, then efficiently covers computational engine, output granularity, key edge cases, sibling differentiation, and idempotency. No filler or repetition; the length is justified by the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 parameters, no output schema, and no annotations, the description is remarkably complete. It tells an agent what the tool does, when to use it, how it behaves (immutability, full position, no posting), what errors to expect (period_locked), and that the return is the draft plus proposed adjusting journal. There are no critical gaps that would block correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the key parameters: asOf/period as cut-off, netRealisableValues as per-unit Veraeusserungswert less costs with the J03 clamp, and inventoryAccountId/changeAccountId as default 1200/4200 overrides. It also explains idempotencyKey via replay semantics. However, workspaceId, method, and notes are not described or hinted at beyond passing mentions, which leaves some parameters underspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb+resource: 'Create a DRAFT inventory valuation run for a cut-off...' and immediately contrasts with inventory_valuation_post ('Posts NO journal'). It also distinguishes the complete-position behavior from inventory_valuation_report for filtered views. An agent can unambiguously identify what this tool does and how it differs from its closest siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it: create drafts for review, with full position valuation; and when not: posting is another tool, filtered views belong to inventory_valuation_report. It also names the exact sibling tools and the conditions that route to them, plus a refusal condition (period_locked) for closed periods. No guesswork is left.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_valuation_getARead-only
One valuation run by id, with its header (as_of, method, status, total value, delta posted, journal link, posted-by) and its frozen lines (item, location, qty, unit cost, value, control account, any OR 960c write-down). A foreign or unknown run id is not_found (H-TENANT), never cross-tenant data.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only convey readOnlyHint=true; the description adds substantial behavior beyond that: frozen-line semantics, exact header components, tenant isolation ('never cross-tenant data'), and the not_found (H-TENANT) error case. This is rich, useful context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence front-loads the core purpose ('One valuation run by id') and packs return structure and error behavior into compact parentheticals. Every clause adds information; there is no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the burden of explaining return values, and it does: header fields, line fields, and the not-found case. For a by-id getter, the description covers parameters, return content, and error semantics completely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does add meaning for runId by tying it to the valuation run and by defining foreign/unknown ID behavior, and workspaceId is implicitly scoped by the tenant-isolation statement. However, it does not explicitly describe workspaceId or give parameter formats/examples, leaving some compensation gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the exact resource ('One valuation run by id') and enumerates the header and frozen-line contents returned. This clearly separates it from list/report/preview inventory valuation siblings, even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It makes clear this is the tool for fetching a single valuation run by its ID, and it signals when it cannot be used by describing the not_found behavior for foreign or unknown IDs. It does not explicitly contrast siblings like inventory_valuation_list or inventory_valuation_report, so it stops short of 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_valuation_layersARead-only
The remaining FIFO cost layers for one item, oldest first: each carries the receipt date, the source movement, the quantity originally received and what is left of it, and its unit cost. Layers are derived from the J02 ledger on demand rather than cached, so they cannot drift from it. Returns the layer quantity, the exact total (the sum of remaining qty times layer cost, to the Rappen) and shortfall, the demand that outlived the layers when stock went short. Filter by locationId to inspect one location; a foreign or unknown item OR location is not_found, never cross-tenant data and never an empty list that would read as "this location holds nothing". NOT A STOCK-AGEING REPORT: when stock is transferred the layer moves with its own unit cost but its receiptDate becomes the date it reached that location, because that is what drives consumption order there. For how long goods have really been held, read the J02 movement history, which rewrites nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| itemId | Yes | ||
| locationId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite readOnlyHint=true covering the safety profile, the description adds rich behavioral context beyond the annotation: layers are derived on demand rather than cached so they cannot drift; the exact total is computed to the Rappen; shortfall semantics are explained; error behavior is disclosed (foreign/unknown item or location is not_found, never cross-tenant data, never an empty list); and the subtle receiptDate rewrite on transfer is explained. This is exemplary disclosure of non-obvious behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but information-dense; every sentence carries a distinct point: scope/ordering, layer fields, derivation freshness, return values, error behavior, and the critical counter-indication. The 'NOT A STOCK-AGEING REPORT' warning is prominently placed. There is minor redundancy between 'derived... on demand rather than cached' and 'which rewrites nothing' at the end, but overall the structure is efficient and well front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lacking an output schema and having 0% parameter coverage, the description covers return values (layer quantity, exact total, shortfall), error semantics, data derivation freshness, and the receiptDate caveat that materially affects interpretation of results. The main shortfall is the unexplained asOf parameter for a tool whose entire purpose is point-in-time layer state. Minor gap, otherwise complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate. It does add meaning for locationId (filter scope and not_found behavior) and implies itemId scope ('for one item'). However, the asOf parameter is never explained — particularly important for an on-demand ledger-derived report — and workspaceId is left implicit. The description partially compensates but leaves a material gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'The remaining FIFO cost layers for one item, oldest first' — stating exactly what is returned, the scope (one item), and the ordering. It further distinguishes itself by explicitly declaring 'NOT A STOCK-AGEING REPORT', making it unmissable that this is not an ageing report and is instead a layer valuation tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use guidance: filter by locationId to inspect one location, and explicitly routes the agent away from this tool when the question is about true holding time ('For how long goods have really been held, read the J02 movement history'). The NOT A STOCK-AGEING REPORT declaration plus the receiptDate mutation caveat tells the agent which alternative to reach for.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_valuation_listARead-only
The valuation runs for this workspace, newest cut-off first: id, as_of, method, status, total value, delta, line count and journal link. Filter by status (draft, posted, reversed), and by an as_of from/to date range. Tenant-scoped; returns headers only (use inventory_valuation_get for the lines).
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| limit | No | ||
| status | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint=true already signals a safe read operation. The description adds behavioral detail: returns headers only, orders by newest cut-off first, and supports filtering by status and as_of date range. No contradiction with annotations, and it provides useful context beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the primary purpose and key fields, followed by filtering and scope details. No wasted words, and it efficiently conveys the necessary information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description lists the returned fields, which is critical. It also explains filtering, scope, and the distinction from the sibling for lines. Minor omissions like pagination or limit behavior are not major for a list tool, so it is fairly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 5 parameters with 0% description coverage, so the description must compensate. It mentions filtering by status and as_of date range, mapping to status and from/to, and implies workspaceId via 'Tenant-scoped'. However, it does not clarify the limit parameter or the exact date format. It partially compensates for the schema gap but leaves some ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists valuation runs for the workspace, enumerates the returned fields (id, as_of, method, status, total value, delta, line count, journal link), and distinguishes it from inventory_valuation_get for lines. This provides a specific verb, resource, and scope, making it easy for an agent to understand its purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly notes the tool is tenant-scoped and returns headers only, with an explicit pointer to inventory_valuation_get for lines. It also mentions filtering by status and date range, giving clear usage context. It doesn't list exclusions (e.g., when not to use), but the alternative is clearly named, so it's strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_valuation_method_historyARead-only
The append-only trail of valuation-method changes, newest effective date first: scope (workspace or item), the method, the standard cost where one was set, the date it took effect, the reason the operator stated, whether it was forced over open-period movements, and who recorded it. This is the OR 958c Stetigkeit evidence: it is what a Treuhänder reads when asked why an inventory figure moved between two years. Filter by itemId or by scope. Rows are immutable at the database layer, so the trail cannot be tidied after the fact.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No | ||
| itemId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds crucial behavioral details: append-only, immutable at the database layer, cannot be tidied, and specific field semantics (e.g., forced over open-period movements). These go beyond the annotation and help the agent understand the guarantees and content.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is informative but not overly verbose. It front-loads the key concept (append-only trail, newest first), lists the returned fields, explains the audit purpose, and adds immutability. Each sentence adds value, though a bit more conciseness could tighten it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description enumerates all returned fields and ordering, and covers immutability and filtering. It lacks pagination details but is otherwise comprehensive for a read-only history tool. The audit context and immutability make it suitable for agents needing to understand the data provenance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explains that filtering is by itemId or scope, and clarifies that scope refers to workspace or item. It does not explicitly describe workspaceId, but it is the required parameter and its role is implied. Given schema description coverage is 0%, the description provides meaningful guidance for two of three parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it returns the append-only trail of valuation-method changes with specific fields (scope, method, standard cost, effective date, reason, forced flag, recorder) and orders by newest effective date first. It names the exact resource and purpose, distinguishing it from current-method or mutation tools like inventory_valuation_methods or inventory_valuation_set_item_method.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains how to filter (by itemId or scope) and provides strong context for when to use it (as OR 958c Stetigkeit evidence for audit). It doesn't explicitly state when not to use it or name alternatives, but the purpose is so specific that the usage is clear from context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_valuation_methodsBRead-only
The valuation method registry: weighted_average, fifo and standard_cost, each with whether this workspace has it enabled and whether it needs a per-item standard cost. Also returns the method in force as the workspace default today, where it came from (a recorded assignment or the built-in weighted_average) and the date it took effect. weighted_average and fifo are enabled out of the box; standard_cost is registered and off until a workspace turns it on. A disabled method can neither be the default nor be chosen on an item.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint=true already covers the read-only nature. The description adds meaningful context beyond that: it explains that the default method's source can be a recorded assignment or the built-in weighted_average, and states that a disabled method cannot be the default or chosen on an item. This provides useful behavioral details without contradicting the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is fairly long but well-structured, covering the registry, default method details, default enablement states, and a constraint. It is front-loaded with the core purpose and uses efficient phrasing, though it could be slightly more concise by trimming some redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does a good job of explaining what data will be returned: method enablement, per-item cost need, default method, its source, and effective date. It also clarifies built-in behavior for weighted_average and fifo. It lacks explicit error handling or response structure details, but for a read-only query this is adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single required parameter workspaceId, and the description does not mention or explain this parameter at all. It fails to compensate for the lack of schema documentation, leaving the agent without any semantic guidance for the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that this tool returns a registry of valuation methods (weighted_average, fifo, standard_cost) with enablement status and per-item cost needs, plus the current default method with its source and effective date. It uses specific verbs and resources, making the purpose clear. It doesn't explicitly differentiate from sibling tools like inventory_valuation_method_history or inventory_valuation_status, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention any sibling tools or conditions for selection. The agent is left to infer that this is for querying the current method registry, but no explicit when/when-not guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_valuation_method_set_enabledA
Turn one valuation method on or off for the workspace. Enablement gates what may be CHOSEN from here on; it is deliberately not dated, because a past figure is determined by the assignment that was in force then and not by what is switched on today. Disabling the method that is currently the workspace default is refused with method_is_default: change the default first, so no future valuation points at a method the workspace says it does not use. Absolute state-setting, so a replay of the idempotency key re-asserts the same list and writes nothing new. Posts no journal entry.
| Name | Required | Description | Default |
|---|---|---|---|
| method | Yes | ||
| enabled | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility for behavioral disclosure. It discloses that the operation is absolute state-setting, idempotent (replay re-asserts the same list, writes nothing new), and posts no journal entry. It also explains the non-dated nature and the error condition method_is_default when disabling the default. This is exceptional transparency for a tool with no annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph of four sentences, each carrying distinct information: primary action, the non-dated rationale, the default-disabling error, and idempotency/no-journal side effects. It is front-loaded with the core purpose. While dense, it avoids redundancy and stays focused, though it could be tightened without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple toggle with four required parameters and no output schema, the description covers the key behavioral aspects: idempotency, error condition, and lack of journal posting. However, it leaves parameter semantics unexplained, which is a gap given 0% schema coverage. An agent may not know what values to pass for 'method' or how to interpret 'enabled'. The description is adequate for when to call it but incomplete for how to fill parameters correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining parameter meanings. It only indirectly references the idempotency key ('replay of the idempotency key') and the enabled flag via 'on or off'. It does not explain what 'method' refers to (name vs. ID), what format workspaceId expects, or the exact semantics of the boolean. The description adds minimal value beyond the schema's field names and types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Turn on or off') and a precise resource ('one valuation method for the workspace'). It immediately clarifies that enablement gates future choices, distinguishing it from valuation methods that set dated assignments or defaults. This clearly separates it from sibling tools like inventory_valuation_set_default and inventory_valuation_set_item_method.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use the tool (to toggle enablement) and gives a critical constraint: disabling the current default is refused, instructing the user to change the default first. It does not explicitly name alternative tools, but the context implies when this tool is the right choice versus others. The guidance is practical and actionable, though it could be more explicit about not using it for dated or default changes.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_valuation_openingA
Record an OPENING inventory valuation for a migration or a new workspace so the sub-ledger = GL identity holds from day one. The caller states the known item values directly (lines of itemId, optional locationId, qty and valueRappen), and the verb posts the opening delta against the current GL through A02 exactly as an ordinary run does, writing a posted run in one step (no draft review: an opening baseline is a stated figure, not a computed one). inventoryAccountId / changeAccountId override the default 1200 / 4200. A cut-off in a locked period is refused before any write. Idempotent on the key.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | Yes | ||
| lines | Yes | ||
| notes | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| changeAccountId | No | ||
| inventoryAccountId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden and does so thoroughly: it posts the opening delta through A02, writes a posted run in one step, refuses locked-period cut-offs before any write, and is idempotent on the key. These are valuable runtime behaviors not visible elsewhere.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary purpose and each subsequent sentence adds distinct operational detail: direct stated values, posting mechanics, account overrides, locked-period protection, and idempotency. There is no filler or repetition of schema content, so every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex write operation with no output schema and no annotations, the description covers purpose, when to use it, exact line semantics, posting behavior, safety protections, and idempotency guarantees. The only omitted details are minor and recoverable from parameter names.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does for the key parameters: lines carry itemId, optional locationId, qty, and valueRappen; inventoryAccountId/changeAccountId override defaults 1200/4200, and idempotencyKey enables idempotence. It leaves asOf and workspaceId semantically implicit, but their field names and the context make them largely inferable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Record an OPENING inventory valuation') and clearly scopes it to migration or new-workspace setups. It also differentiates itself from ordinary valuation runs by stating it posts directly with no draft review and that values are stated rather than computed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear when-to-use context: migration/new workspace and when the caller has known item values to establish sub-ledger = GL identity from day one. It implies when not to use it (computed or draft-review scenarios) by contrasting with 'an ordinary run', though it does not explicitly name an alternative sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_valuation_postA
Post a reviewed DRAFT run so the balance-sheet inventory figure is updated in one atomic, auditable step. In one transaction it re-checks the cut-off period is open and the draft is not stale (the movement ledger has not changed since it was calculated, else stale_draft), computes the delta between the sub-ledger valuation and the current GL balance of each control account, posts ONE balanced A02 entry (source inventory_valuation, Dr/Cr inventory control vs the change account) for that delta, and marks the run posted with its journal reference. After the post the OP11 identity holds: GL inventory control balance equals the sub-ledger valuation at the cut-off. All posting is through A02 postEntry and no second path exists. Idempotent on rows: a replay returns the posted run and mints no second journal. A zero delta posts nothing yet records the baseline. CONSEQUENCE: Posts the inventory valuation delta and updates the balance-sheet inventory figure.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| changeAccountId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and excels. It discloses atomicity, re-checks (cut-off period, stale draft), delta computation, the A02 entry posting, idempotency, zero-delta behavior, and the resulting OP11 identity. It even states that no second posting path exists. This is exceptionally transparent about side effects and behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-organized. It front-loads the purpose and consequence, then details the process, idempotency, and edge cases. It is longer than minimal but every sentence adds value—no fluff. The structure (purpose, process, idempotency, consequence) makes it easy to scan, though it could be slightly more concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex mutation tool with no output schema, the description covers the critical behavioral context: atomicity, checks, idempotency, and the resulting GL balance. It does not explicitly describe the response format (e.g., journal reference returned) or error scenarios beyond stale_draft, but the mention of 'marks the run posted with its journal reference' implies what to expect. Overall, it is quite complete for the agent to call correctly, though a brief note on return values would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% coverage, so the description must explain each parameter. It explicitly mentions idempotencyKey and the 'change account' (changeAccountId) in the context of the delta posting, and implicitly refers to runId as the 'draft run'. However, workspaceId is not mentioned, and changeAccountId is only alluded to as 'the change account' without explicitly mapping it to the schema field. The description provides context but does not give a clear parameter-by-parameter explanation, which is a gap given the 0% schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Post'), a precise resource ('a reviewed DRAFT run'), and the exact outcome (updates the balance-sheet inventory figure). It clearly distinguishes from sibling tools like inventory_valuation_preview and inventory_valuation_reverse by describing the posting action and its atomic, auditable nature. No ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: after a draft run has been reviewed, and it is the posting step in the valuation workflow. It does not explicitly name alternatives or state 'use this instead of preview' but the context of 'post' vs 'preview' and the mention of 'reviewed DRAFT run' make the usage context clear. However, it lacks an explicit exclusion like 'do not use if the draft is not reviewed' or a comparison to inventory_valuation_reverse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_valuation_previewARead-only
Value inventory at a date WITHOUT writing anything: no journal entry, no run row, no cached figure. Returns one row per item with the method that applied on that date (item override beats workspace default beats the built-in weighted_average, each as it stood on the asOf date, so a later policy change never restates a filed figure), qtyOnHand, the rounded unitCostMinor a human reads, and totalValueMinor computed from the exact cost pool. Omit itemIds to value every item with ledger history; omit asOf for everything recorded so far. methodOverride projects a what-if under another method and leaves the stored policy untouched (a disabled method is refused). netRealisableValues maps itemId to the OR 960c per-unit Veräusserungswert LESS the costs still to come; where it is below cost the clamp is applied, lcmApplied is set and writeDownMinor reports the difference. Uncosted receipts are valued at zero, counted in uncostedQty and named in missingCostMovementIds, never valued at the neighbours price. An internal transfer between locations is value-neutral: cost follows the goods. LOCATION SCOPE: locationId alone values that one location, valueByLocation alone returns one row per location, both together narrow the breakdown; an unknown or foreign locationId is not_found, never an empty result. Per-location rows always sum to the item total, and each says how it got there in valuationBasis: direct for FIFO and standard cost, allocated for weighted average, which has one cost pool per item and so shows each location its own quantity at the item pooled average rather than inventing a second pool. A negative on-hand, a negative unit cost or a missing standard cost returns a reason and a value of zero rather than an arithmetic answer.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| itemIds | No | ||
| locationId | No | ||
| workspaceId | Yes | ||
| methodOverride | No | ||
| valueByLocation | No | ||
| netRealisableValues | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description reveals many behavioral traits: no cache or persistence, method precedence rules, date-dependent policy snapshots, rounding of unit costs, handling of uncosted receipts and negative balances, location aggregation with valuationBasis, and edge-case errors. This is extensive and valuable transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long (about 250 words) but every sentence adds substantive detail about behavior, parameters, or edge cases. It is front-loaded with the core purpose, then organized into logical sections (calculation details, location scope, error conditions). Minor redundancy exists (e.g., internal transfer sentence could be shorter), but it is not bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (7 params, nested objects, no output schema), the description covers output fields, edge conditions, error handling, and all parameters. It also explains how per-location rows sum to item totals and what valuationBasis values mean. Nothing critical is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate — and it does. It explains itemIds (omit for all), asOf (omit for all recorded), methodOverride (what-if projection, disabled methods refused), netRealisableValues (mapping and clamp logic), locationId vs valueByLocation (scoping and output shape), and implicitly the required workspaceId. No parameter left unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Value) and resource (inventory at a date) and explicitly contrasts itself with write operations: 'WITHOUT writing anything: no journal entry, no run row, no cached figure.' This distinguishes it from sibling tools like inventory_valuation_post or stock_run_valuation as a read-only preview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use it: when you need to value inventory without posting, as emphasized by 'WITHOUT writing anything.' It does not explicitly name alternatives or say 'use X instead,' but the read-only framing and details like 'methodOverride ... leaves the stored policy untouched' provide clear context for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_valuation_reportARead-only
The authoritative detailed valuation report: the single source of truth for what inventory is worth at a cut-off. With runId it returns the FROZEN lines of that posted run; otherwise it computes the LIVE valuation at asOf (or period end) through the J03 engine, grouped by control account, with each line item, location, quantity, unit cost and value plus any OR 960c write-down. Optional accountIds narrows the account grouping; method projects a what-if under another method; netRealisableValues applies the lower-of-cost-or-market clamp. Writes nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| runId | No | ||
| method | No | ||
| period | No | ||
| groupBy | No | ||
| accountIds | No | ||
| workspaceId | Yes | ||
| netRealisableValues | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces this with 'Writes nothing.' It adds behavioral detail beyond annotations by explaining the dual modes (frozen vs live), the use of the J03 engine, grouping by control account, and the lower-of-cost-or-market clamp. This provides valuable context without contradicting the read-only hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense paragraph that front-loads the core purpose and then methodically covers conditional behavior, grouping details, and optional parameters. Every sentence adds value with no fluff. The 'Writes nothing' ending is a clear, concise closure. It is appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the 8 parameters, nested objects, and no output schema, the description covers the essential aspects: what the report returns (line items, location, quantity, unit cost, value, write-downs), how it computes (J03 engine), the two modes, and optional parameters. It does not explain the groupBy parameter or potential pagination/limits, but these are minor gaps. Overall it provides sufficient context for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden of explaining parameters. It explains runId (frozen lines), asOf (live valuation date), period (period end), accountIds (narrow account grouping), method (what-if under another method), and netRealisableValues (applies lower-of-cost-or-market clamp). It does not explain groupBy or workspaceId, but these are relatively self-explanatory. This compensates well for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it is the authoritative detailed valuation report, the single source of truth for inventory worth at a cut-off. It specifies it returns frozen lines for a posted run or computes live valuation via J03 engine, with explicit detail on grouping and optional parameters. This distinguishes it from sibling reports like stock_valuation_report or inventory_valuation_preview, which are implied to be less authoritative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear conditional usage: with runId it returns frozen lines, otherwise it computes live valuation at asOf or period end. It also explains optional parameters like method for what-if and netRealisableValues for lower-of-cost-or-market. However, it does not explicitly name alternatives or state when not to use this tool, though the 'single source of truth' phrasing implies it should be preferred over siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_valuation_reverseA
Reverse a posted valuation run. Posts the exact A02 mirror of the run journal via reverseEntry (which moves the GL back by the delta and restores the baseline every later run measures against), marks the original run reversed and records the reversing journal. The original run and its journal are never mutated: a correction is a reverse plus a fresh run. A reason is required. Idempotent: a replay returns the already-reversed run and posts no second compensation. A run that posted no journal (zero delta) is marked reversed with nothing to reverse. CONSEQUENCE: Posts the mirror of a valuation run and restores the baseline the next run measures against.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| reason | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure and does so exceptionally. It reveals side effects (GL moves back by delta, baseline restored), what is not mutated (original run and journal), idempotency replay behavior, the zero-delta edge case, and the requirement of a reason. It even labels the consequence explicitly. This is far beyond typical descriptions and leaves no major behavioral surprises.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary action and then provides dense, useful details. It is structured to flow from what it does to edge cases. However, there is redundancy: the opening states the mirror journal and the 'CONSEQUENCE' line repeats essentially the same information. Trimming that duplication would make it even tighter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, no output schema, and no annotations, the description covers a remarkable amount of behavioral context: side effects, idempotency, mutation policy, and the zero-delta case. The main missing piece is parameter-level guidance (especially idempotencyKey semantics), which prevents a perfect score. Overall, an agent could invoke this safely but might still be unsure how to construct the arguments.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate for parameter documentation. It only mentions that a 'reason is required' but does not explain the meaning or expected format of runId, workspaceId, or idempotencyKey. The reader cannot infer how to populate these parameters beyond their names, leaving a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Reverse a posted valuation run.' It then details the exact mechanism (posts A02 mirror via reverseEntry), marking the run reversed, and explicitly states the original run is never mutated, which distinguishes it from other reversal tools in the sibling list (e.g., asset_depreciation_run_reverse, fx_revaluation_reverse). This makes the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly explains the context in which this tool is used: for reversing a posted valuation run and how corrections are structured ('a correction is a reverse plus a fresh run'). It also notes the idempotent behavior and the zero-delta case. However, it does not explicitly name alternative tools or state when not to use this tool, such as using reverse_entry for general journal reversals, so it falls slightly short of an explicit when/when-not guide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_valuation_set_defaultA
Set the workspace default valuation method from effectiveFrom on. Appends an immutable assignment row rather than overwriting a setting, so the history of what applied when survives (OR 958c Stetigkeit) and a valuation at an earlier date still resolves to the method that was in force then. effectiveFrom is what the period lock is checked against, NOT the day you call it: a change dated into a soft- or hard-closed period is refused with period_locked before anything is written, so a sealed year cannot be restated by picking an old date. When the workspace already has movements dated on or after effectiveFrom, the change would restate their valuation for every item without an override of its own, so it is refused with method_change_blocked_open_period unless you pass forceRevaluation true AND a reason (a blank reason does not count). A disabled method is refused. Posts no journal entry.
| Name | Required | Description | Default |
|---|---|---|---|
| method | Yes | ||
| reason | No | ||
| workspaceId | Yes | ||
| effectiveFrom | Yes | ||
| idempotencyKey | Yes | ||
| forceRevaluation | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and exceeds it: it explains the immutable assignment row, how effectiveFrom interacts with period locks, the method_change_blocked_open_period condition, forceRevaluation and reason requirements, disabled-method refusal, and the fact that no journal entry is posted. This is far beyond what the schema reveals.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but front-loaded with the primary action, then each subsequent sentence covers a distinct failure mode or side effect without repeating schema boilerplate. The OR 958c Stetigkeit parenthetical and dense wording are minor readability costs, but no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a complex mutation with hidden side effects, no annotations, and no output schema, yet the description covers temporal resolution, period-lock semantics, open-period blocking, forced revaluation prerequisites, disabled-method rejection, and absence of journal impact. That is sufficient for an agent to call it correctly in the main scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does for key parameters: effectiveFrom gets extensive semantic treatment, forceRevaluation and reason are explained as jointly required, and method is clarified via disabled-method refusal. workspaceId and idempotencyKey remain conventional and are not explicitly described, but they are less domain-critical.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource: 'Set the workspace default valuation method from effectiveFrom on.' It further distinguishes this workspace-default operation from per-item behavior by referencing 'every item without an override of its own,' which differentiates it from sibling tools like inventory_valuation_set_item_method.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear contextual guardrails: soft- or hard-closed periods cause period_locked refusal, existing movements require forceRevaluation=true plus a non-blank reason, and disabled methods are refused. It does not explicitly name alternative tools such as inventory_valuation_set_item_method as an alternative for item-level changes, so it stops short of full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_valuation_set_item_methodA
Override the valuation method for one item from effectiveFrom on, with standardCostMinor when the method is standard_cost (required, integer Rappen, above zero). Appends an immutable assignment row; effectiveFrom is what the period lock answers to, never the call date. When the item already carries movements dated on or after effectiveFrom, those are exactly the movements this change restates, so it is refused with method_change_blocked_open_period unless you pass forceRevaluation true AND a reason: the reason is the sentence that appears in the Stetigkeit history when someone asks why the figure moved, so a force without one is refused rather than recorded blank. The revaluation itself belongs to J06. A disabled method is refused; a foreign item is not_found. Posts no journal entry.
| Name | Required | Description | Default |
|---|---|---|---|
| itemId | Yes | ||
| method | Yes | ||
| reason | No | ||
| workspaceId | Yes | ||
| effectiveFrom | Yes | ||
| idempotencyKey | Yes | ||
| forceRevaluation | No | ||
| standardCostMinor | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden and does exceptionally: appends immutable row, effectiveFrom is the lock key, revaluation belongs to J06, posts no journal entry, force requires reason for history transparency. Contradicts nothing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and information-rich but a single block of text; could be structured with bullet points or separated into conditions. However, every sentence adds critical behavior information, and key points are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, description covers failure modes, side effects (no journal entry), history implications, and conditions for bypass. An agent has enough to decide when to call and what to provide; missing only explicit value formats for dates, but schema holds type hints.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but description explains standardCostMinor required for standard_cost, effectiveFrom semantics (lock answer, not call date), forceRevaluation and reason interplay, and idempotencyKey implied by 'immutable assignment row'. Provides essential meaning beyond bare schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Override'), resource ('valuation method for one item'), and scope ('from effectiveFrom on'), and specifies the required standardCostMinor for standard_cost. It clearly distinguishes from related tools like inventory_valuation_set_default by focusing on per-item override rather than default setting.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear conditions: when movements exist on/after effectiveFrom, refusal unless forceRevaluation and reason; disabled methods refused; foreign items not_found. Does not explicitly name alternative tools beyond mentioning the method_change_blocked_open_period error, but the context is sufficient for typical use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
inventory_valuation_statusARead-only
Der Bewertungsstatus (valuation status / drift, US-J07.3, the OP11 surface): recompute the current inventory value LIVE from the J02 ledger through the J03 method at as_of (default today, optional method override), read the most recent POSTED J06 valuation run at or before that cut-off, and report current_value_rappen, last_posted_value_rappen, drift_rappen, drift_pct, last_posted_at, last_run_id and a status of aligned | drift_present | never_posted. Drift is REPORTED, never auto-corrected: J07 posts nothing and reaches no GL. With include_detail the per-item current values are returned. This is the hard check the period-close checklist reads. Zero inventory answers zeros, aligned.
| Name | Required | Description | Default |
|---|---|---|---|
| as_of | No | ||
| method | No | ||
| thresholds | No | ||
| workspaceId | Yes | ||
| include_detail | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds important behavioral context beyond the readOnlyHint annotation: 'Drift is REPORTED, never auto-corrected: J07 posts nothing and reaches no GL.' It also clarifies that the computation is live and that zero inventory yields aligned status. These are valuable details about side effects and edge cases. It does not contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured, front-loading the core purpose and the key status list. It avoids redundancy and includes only necessary details like the GL impact and the zero-inventory edge case. Some internal jargon ('OP11 surface', 'J02', 'J03') is used without explanation, but it adds precision for knowledgeable users.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, the description enumerates all returned fields explicitly. It also covers edge cases (never_posted, zero stock) and side effects. The main gap is the undefined thresholds parameter, but the tool's core function and expected output are sufficiently specified for an agent to call it correctly in typical scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains as_of (default today), method override, and include_detail (per-item values). However, it does not explain thresholds or the required workspaceId. It covers three of five parameters clearly, leaving two unspecified. This partial coverage warrants a mid-level score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states exactly what the tool does: recomputes current inventory value live from the J02 ledger via the J03 method, compares it to the most recent posted J06 valuation, and reports drift with a status. The verb 'recompute' and the resource 'inventory valuation status' are specific channels. It clearly differentiates itself from other inventory tools by returning a status with drift and posting nothing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: 'This is the hard check the period-close checklist reads.' It implies this is a read-only validation tool versus other tools. However, it does not explicitly name alternative tools or define when not to use it, leaving some inference to the agent. Overall, usage context is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
invite_memberA
Invite someone to this workspace with a role. Writes a pending membership and PREPARES the invite; TILL has no transport and never claims a send, so the token is handed back for the operator to deliver. kind (human, the default, or agent) says what the invitee is: an agent member is the governed seat, its writes draft under the approval dial and its calls land in the trace, on every door including a served one. CONSEQUENCE: Seats a person or an agent in the workspace with the chosen role; whoever redeems the token sees the books from then on.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| role | Yes | ||
| Yes | |||
| displayName | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full disclosure, and it is strong: it states the write side effect ('Writes a pending membership'), what it does not do ('never claims a send'), and the long-term consequence ('sees the books from then on'). The agent-kind explanation adds meaningful behavioral context about governance and approval/tracing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary action, and each major clause adds behavioral or consequence information rather than repeating the name. It is dense but contains product-specific jargon ('governed seat', 'served one') that an operator may have to decode, so it is not maximally crisp.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and six parameters, the description covers the tool's side effects and no-send behavior well. However, it omits the return shape of the token, role value discovery, idempotency semantics, and any permissions prerequisites, so an agent still has to consult sibling tools or conventions for full safe use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema descriptions cover 0% of the six parameters, so the description must compensate. It does meaningfully explain the only non-obvious parameter, kind (human default vs. agent). But it never clarifies idempotencyKey's role or any allowed role values, and displayName is left entirely to the schema, so several parameters remain under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening clause states a specific verb and resource ('Invite someone to this workspace with a role'), and the rest clarifies that it only prepares a pending invitation rather than sending it. This makes it clearly distinct from accept_invite and set_role in the sibling set. The consequence statement reinforces the intended effect.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to call it – when a pending membership and a delivery token are needed, and explicitly warns that TILL will not send the invite. It does not, however, name alternatives or state when not to use it (e.g., accept_invite for redemption, set_role for changing an existing member).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
issue_credit_noteB
Issue a drafted Gutschrift: assign the gap-free G-number and post the mirror entry (debit revenue and output VAT, credit 1100 Debitoren) dated today, closing against the invoice per rate class (D67/D71). CONSEQUENCE: Freezes the credit note, assigns its number and posts its ledger effect.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| creditNoteId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden and does well: it states the consequence up front ('Freezes the credit note, assigns its number and posts its ledger effect'), the exact posting accounts, and the fact that it is dated today. It does not mention reversibility or error conditions, but the 'freezes' wording strongly signals irreversibility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, but the CONSEQUENCE sentence largely repeats the first sentence's content ('assigns its number and posts its ledger effect'). The repetition costs it a higher score; some of the operational jargon (Gutschrift, G-number, D67/D71) could be clarified without adding length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description provides substantial operational context: posting accounts, date behavior, invoice closing, and the freezing consequence. But it omits parameter semantics, preconditions, what constitutes 'gap-free,' and what the caller should expect as a return value, leaving meaningful gaps for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions workspaceId, creditNoteId, or idempotencyKey. An agent must infer what the creditNoteId refers to and why an idempotencyKey is required. The description fails to compensate for the schema's lack of semantic documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Issue') on a specific resource ('a drafted Gutschrift' / credit note), and immediately distinguishes it from the sibling create_credit_note by emphasizing it acts on an already-drafted document. The operational details (assign G-number, post mirror entry, close against invoice) make the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage context is implied by the phrase 'a drafted Gutschrift' — the tool is for issuing an existing draft, not creating a credit note. However, it never explicitly states when to use this tool versus create_credit_note, issue_invoice, or send_invoice, nor does it state preconditions like 'credit note must be in draft state' or 'invoice for the period must be open.'
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
issue_dunning_runA
Issue a proposed Mahnlauf: re-check every item against the OP-Liste as of today (a settled invoice drops out, a shrunken one shrinks) AND against the escalation state (one issued reminder per level per invoice: an item another run already issued at this level drops as stale), freeze the items and the Verzugszins note figures (Art. 104 OR, on the invoice principal), and, when the policy books a Mahngebühr, post ONE entry (debit 1100 Debitoren, credit the fee-income account, the fee's VAT split pro rata across each chased invoice's own rate bases per D69). A locked period skips the fee, names it on the run, and the run still issues; calling issue again on that run with a NEW idempotency key BOOKS the skipped fee once the period is open, and the booking never rewrites the letter (D73: the demand froze at issue; the recovered fee joins the next escalation letter or ordinary collection on the OP-Liste). Reusing the ORIGINAL issue key on such a run replays the issue's own answer while the period is still locked, and refuses recovery_needs_its_own_key once it is open (a retry and a recovery are indistinguishable under one key, so the engine refuses rather than guesses; the payload names the remedy). P8: pass confirmed=true or enable the workspace dial. The letters exist from here on via get_dunning_pdf. A mistaken fee is corrected by reverse_entry, never an edit. CONSEQUENCE: Freezes the reminder run, advances each debtor one dunning level and books any fee.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| confirmed | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does an excellent job disclosing side effects: freezing the reminder run, advancing debtors one dunning level, booking a fee, and the locked-period behavior. It also discloses a consequence (CONSEQUENCE section) and invariants (fee correction via reverse_entry, letters via get_dunning_pdf). The only minor gap is that it doesn't explicitly state permissions/auth requirements or rate limits, but for a domain operation with no annotations, this is remarkably transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and front-loaded with the core action, but it is extremely long and packed with parentheticals and legal/domain jargon (D69, D73, Art. 104 OR, OP-Liste, Mahnlauf). Many sentences could be tightened, and the dense clauses ('the fee's VAT split pro rata across each chased invoice's own rate bases per D69') may obscure rather than clarify for an AI agent. Still, every sentence carries real semantic content and there is no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's high complexity (idempotency nuances, locked-period behavior, fee posting, VAT split) and the absence of both annotations and an output schema, the description is remarkably complete. It explains the full lifecycle, the consequence, the correction path, and the idempotency contract. The only missing items are explicit return value details, but with no output schema the description still gives enough context for correct invocation. It is complete enough to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description names only two of the four parameters (confirmed and idempotencyKey), but it gives deep semantics for those two: confirmed=true or workspace dial enables the action, and idempotencyKey has distinct behavior for new vs original keys under locked periods. It does not explicitly explain runId or workspaceId, but their purpose is strongly implied by the domain context. Given the 0% schema coverage, the description compensates substantially for the key behavioral parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Issue a proposed Mahnlauf') and exhaustively details what the tool does: re-checking items against the OP-Liste and escalation state, freezing items and Verzugszins figures, posting one Mahngebühr entry, and advancing dunning levels. It clearly distinguishes this from siblings like propose_dunning_run, get_dunning_pdf, and send_dunning_run, making the action unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly explains when to issue (after a proposal, with confirmed=true or the workspace dial enabled) and what happens under edge cases: locked periods skip the fee, calling with a NEW idempotency key books the fee later, reusing the ORIGINAL key replays the issue's answer, and retry vs recovery is refused. It also tells the agent to use reverse_entry for corrections, not edit. This is explicit when/when-not guidance and names alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
issue_invoiceA
Issue a draft invoice: assign the gap-free number, post the balanced VAT entry (A10 onIssue delegate), and attach the QR-bill summary. CONSEQUENCE: Freezes the invoice and assigns its number; an issued invoice can only be credited, never edited.
| Name | Required | Description | Default |
|---|---|---|---|
| invoiceId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well: it names the side effects (number assignment, VAT posting, QR attachment) and the irreversible consequence (freeze, credit-only). It could also mention prerequisites or failure behavior, but the explicit 'CONSEQUENCE' framing is strong for an issue operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, with the action front-loaded and the consequence clearly labeled. The parenthetical '(A10 onIssue delegate)' is dense and possibly confusing, but it does not prevent comprehension of the main behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two required scalar parameters and no output schema, this description is largely complete: it states what happens, what gets frozen, and what the long-term constraint is. It leaves room for error/precondition details and idempotencyKey semantics, but the core call is understandable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description does not explain workspaceId, invoiceId, or idempotencyKey. invoiceId and workspaceId are somewhat self-evident from their names, but idempotencyKey and the additionalProperties behavior are left entirely undefined, and the description does not compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Issue a draft invoice') and enumerates the precise mechanics: assign the gap-free number, post the balanced VAT entry, and attach the QR-bill summary. The draft-vs-issued distinction and the 'credited, never edited' consequence separate it from sibling tools like send_invoice and issue_credit_note.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly indicates the tool is for draft invoices that are ready to be frozen and numbered. The warning that an issued invoice can only be credited, never edited, gives a when-not boundary. However, it does not explicitly name alternatives like create_credit_note or send_invoice, so routing guidance is inferred rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
item_categories_deleteA
Delete an item category. Refused (category_in_use) while any item or child category still points at it; reassign or empty it first.
| Name | Required | Description | Default |
|---|---|---|---|
| categoryId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the refusal condition (category_in_use) and the prerequisite, which is valuable. However, it does not mention other behavioral aspects such as idempotency, whether the deletion is permanent, or any side effects beyond the failure mode. The description adds some context but not comprehensive transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero fluff. It leads with the action, then immediately states the critical failure condition and remedy. Every word earns its place, making it easy to parse and act on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple delete operation, the description covers the core purpose and failure mode, but it omits parameter semantics entirely and does not describe the return value or confirmation of success. Since there is no output schema and no parameter descriptions, an agent may be uncertain about required inputs and what to expect on success. The description is too thin for a tool with three parameters and no annotation support.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the lack of parameter documentation. It fails to do so: it never mentions categoryId, workspaceId, or idempotencyKey, nor does it explain their roles. The agent must infer that categoryId identifies the category and workspaceId scopes the operation, which is not evident from the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Delete an item category.' It uses a specific verb and resource, and it distinguishes this tool from its siblings (item_categories_upsert, item_categories_list) by its delete function. There is no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear precondition: it will be refused while any item or child category still points at it, and advises to reassign or empty first. This tells the agent when the tool will fail and what to do. However, it does not explicitly contrast with alternative delete tools for other resources, but given the resource is unique, it's sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
item_categories_listARead-only
List the item categories in the workspace, ordered by sort then name.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers the safety profile. The description adds a genuine behavioral detail beyond that – result ordering ('by sort then name') – which is useful. However, it does not disclose pagination, whether archived categories are included, or the shape of the returned items, so coverage is decent but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that is entirely front-loaded: verb, resource, scope, and ordering are stated with zero filler or redundancy. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with one parameter and no output schema, the description is nearly complete: it names the resource, scope, and ordering. Minor gaps (pagination, archived-item inclusion, return shape) are acceptable given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the undocumented workspaceId parameter. It partially does by scoping the resource to 'the workspace', implying workspaceId selects the workspace. But it adds no detail about the parameter's format, source, or default behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('list'), a precise resource ('item categories'), scope ('in the workspace'), and even the ordering ('by sort then name'). This clearly distinguishes it from the mutation siblings item_categories_upsert and item_categories_delete, and from generic list tools like list_items.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not-to-use guidance is provided. The read intent is inferable from the verb 'list' and the sibling names (upsert/delete), but the description itself never states when this tool should be chosen over alternatives or mentions exclusions such as archived categories.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
item_categories_upsertA
Create or edit an item category (two-level tree: the parent of a child must itself be a root). Pass categoryId to edit, omit it to create.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| sort | No | ||
| parentId | No | ||
| categoryId | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden. It does reveal a key invariant (two-level tree: parent of a child must be root), which adds useful context. However, it does not disclose potential side effects when editing a category that has children, nor any permissions or reversibility requirements. The mutation intent is clear but not deeply transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence packs the key information: upsert semantics, tree hierarchy constraint, and the categoryId conditional. No filler or repetition; the sentence is front-loaded with the action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite the clear purpose and tree constraint, the description omits crucial details for correct invocation: the required workspaceId is never mentioned, and parameter semantics for name, sort, parentId, and idempotencyKey are absent. Given no output schema, an agent would need to infer too much from parameter names alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for six undocumented parameters. It explains the role of categoryId (edit vs create) but leaves the meaning of name, sort, parentId, workspaceId, and idempotencyKey entirely to schema property names. This is a significant gap for an agent trying to invoke correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb ('Create or edit') and resource ('item category') and clearly differentiates the two modes via categoryId. It also distinguishes from sibling tools like item_categories_delete and item_categories_list, making the tool's role unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly explains when to edit (pass categoryId) versus create (omit categoryId). It does not name alternatives or exclusions, but the upsert pattern is clear and the two-mode guidance is actionable without further inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
item_set_tracking_modeA
Set an item's lot/serial tracking mode: none | lot | serial | lot_and_serial. The item must be stockable (track_stock true) and its current on-hand must be 0 for any non-none mode (else tracking_mode_requires_zero_stock); a service or non-stockable item is tracking_not_applicable. Going back to none needs every lot and serial for the item archived and balanced. Plain master data: posts no journal entry.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | Yes | ||
| itemId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full disclosure burden. It states that this is plain master data and posts no journal entry, and it reveals the constraint-heavy consequences of switching modes. This goes well beyond what the bare schema or annotations would provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient: it front-loads the operation and mode values, then gives preconditions/errors, revert requirements, and side-effect disclosure in three short follow-up sentences. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description covers purpose, valid modes, prerequisites, error codes, and the no-journal-entry side effect. The main omissions are explicit semantics for idempotencyKey and what a successful response contains, but these are not critical for invoking the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description is the only source of parameter meaning. It fully documents the mode parameter with exact allowed values and links it to item state. itemId is naturally implied, but workspaceId and idempotencyKey are not explicitly explained, leaving a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Set an item's lot/serial tracking mode,' followed by the four legal values. This clearly distinguishes it from sibling inventory/item operations like serial_set_status or update_item, which target different concerns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage conditions: the item must be stockable (track_stock true), on-hand must be 0 for non-none modes, service/non-stockable items are rejected, and reverting to none requires all lots/serials archived and balanced. It even names the relevant error codes, letting an agent predict failure before calling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
landed_cost_allocate_confirmA
Confirm the allocation of a DRAFT voucher, in ONE atomic transaction: re-run the pure allocator (verifying the residual is 0), write one value-only J02 landed_cost movement per target (qty 0, the allocated Rappen as cost_amount, ref_movement_id = the receipt movement, so J03 folds the cost onto that exact layer and on-hand is unchanged), post ONE balanced A02 entry, read it back to verify it balances, and move the voucher draft -> allocated. The GL entry SPLITS the cost by what J03 actually carries: Dr inventory control for the CAPITALIZABLE share (the on-hand fraction for weighted-average and FIFO, ZERO for a standard-cost item), Dr the variance account for the remainder (already-issued units, and a standard-cost item whole cost), Cr landed-cost clearing for the total, so the inventory-control debit equals the J03 valuation rise exactly (OP11). variancePolicy strict refuses a remainder (landed_cost_variance_not_permitted); expense_excess needs varianceAccountId set (variance_account_required). method and effectiveDate default to the voucher. A replay under the same idempotencyKey returns the allocated voucher and writes nothing; a second confirm under a different key is invalid_transition. period_locked when the effective date is in a sealed period. CONSEQUENCE: Allocates the voucher across its targets and posts the balanced landed-cost entry.
| Name | Required | Description | Default |
|---|---|---|---|
| method | No | ||
| voucherId | Yes | ||
| workspaceId | Yes | ||
| manualShares | No | ||
| effectiveDate | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it excels: it discloses the atomic transaction steps, the exact movements written, the balanced A02 entry, idempotent replay, invalid transitions, and period-lock errors. The GL split explanation and variancePolicy consequences go far beyond typical tool descriptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place, covering workflow, accounting mechanics, error cases, defaults, and idempotency. The colon-led structure front-loads the core action and the final CONSEQUENCE line reinforces the outcome without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a multi-step allocation, GL posting, and idempotency semantics, the description is remarkably complete: it documents the full sequence, expected invariants, error conditions, defaults, and replay behavior. Missing return-value details are acceptable since no output schema exists and the primary outcome is clearly stated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning to method and effectiveDate (defaults), and idempotencyKey (replay vs. invalid transition), but manualShares is a nontrivial object parameter left unexplained, and workspaceId/voucherId are only implied by their names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description precisely states a specific verb and resource: 'Confirm the allocation of a DRAFT voucher' and details the atomic transition from draft to allocated. It clearly differentiates this from siblings like landed_cost_allocate_preview and landed_cost_reverse by describing the commit action rather than a preview or reversal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear context on when to use the tool: only for DRAFT vouchers, with replay behavior and defaulting rules. It does not explicitly name alternatives like landed_cost_allocate_preview or state 'use preview before confirming,' so it stops short of explicit when-not/exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
landed_cost_allocate_previewARead-only
PURE allocation preview of a voucher: per target the base value and quantity, the share (0..1), the allocated Rappen and the per-unit impact, the effective method actually used (a fallback swaps by_weight / by_volume to by_value when an attribute is missing, reported in warnings), and the residual (always 0 after commercial rounding). method defaults to the voucher method; manualShares supplies per-target fractions for the manual method (must sum to 1). Writes nothing and is safe to call repeatedly; changing the method immediately re-computes.
| Name | Required | Description | Default |
|---|---|---|---|
| method | No | ||
| voucherId | Yes | ||
| workspaceId | Yes | ||
| manualShares | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses important behavior: fallback logic from by_weight/by_volume to by_value, warning reporting, residual always being 0 after rounding, immediate re-computation on method change, and safety of repeated calls. This is substantial value beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause earns its place: outputs, fallback semantics, defaults, parameter constraints, and safety. It front-loads the core purpose before moving into details and contains no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description thoroughly covers what the tool returns, how the method is resolved, what fallback behavior occurs, and what can be configured. An agent has enough information to call and interpret this preview correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates well: it explains that method defaults to the voucher method and names possible values via the fallback discussion, and it defines manualShares as per-target fractions for the manual method with the sum-to-1 constraint. workspaceId and voucherId are not discussed, but their meaning is self-evident from schema names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'PURE allocation preview of a voucher' and enumerates the concrete output fields (base value, quantity, share, allocated Rappen, per-unit impact, method, residual). This clearly distinguishes it from the sibling landed_cost_allocate_confirm, which performs the actual allocation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'preview' framing and 'Writes nothing and is safe to call repeatedly' make the intended simulation use-case clear. It does not explicitly name the alternative confirm tool or state 'use confirm to apply the allocation', so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
landed_cost_getBRead-only
One landed-cost voucher with its header, cost-component lines and allocation targets (each with the goods-receipt line, the original receipt movement, the allocated amount and unit impact, and the J02 movement / reversal movement it minted). A foreign or unknown id is not_found, never cross-tenant data.
| Name | Required | Description | Default |
|---|---|---|---|
| voucherId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, and the description adds valuable behavioral context: the not_found result for foreign/unknown ids and the guarantee of data isolation across tenants. It also enumerates the detailed response contents, which goes beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the main object and then details its components. It is efficient with no filler, though the nested parenthetical is somewhat complex.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates well by detailing the return structure and covering error cases. It remains incomplete on parameter semantics and usage context, but for a simple getter, the core calling information is mostly present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero descriptions for workspaceId and voucherId. The description mentions 'a foreign or unknown id' but does not explicitly map this to voucherId, nor does it clarify workspaceId semantics beyond its name. Some meaning is added, but insufficiently given the 0% schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource: a single landed-cost voucher with its header, cost-component lines, and allocation targets. The verb 'get' in the tool name plus the focus on one voucher make the purpose evident, though it does not explicitly contrast with the list sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives like landed_cost_list or landed_cost_voucher_create. The description is entirely about output structure and error behavior, not about selection criteria or preconditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
landed_cost_listARead-only
List landed-cost vouchers with filters: status (one value or an array of draft | allocated | reversed), itemId (vouchers that touch that item), and an effectiveDate range (fromDate / toDate). Each row carries the number, status, estimated flag, currency, total cost in Rappen, allocation method, effective date and the linked journal entry id. Newest effective date first. A foreign workspace sees only its own vouchers (§H-TENANT).
| Name | Required | Description | Default |
|---|---|---|---|
| itemId | No | ||
| status | No | ||
| toDate | No | ||
| fromDate | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses the sort order ('Newest effective date first'), tenant isolation behavior (§H-TENANT), and the exact row content (number, status, estimated flag, currency, total cost in Rappen, allocation method, effective date, linked journal entry id). No contradiction with annotations: 'List' aligns with readOnlyHint=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, no filler: filters, row content, sort order, and tenant scoping each get exactly the space they need. The core action is front-loaded and every clause contributes a distinct fact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema present, the description compensates by enumerating row fields, and it specifies ordering, filtering, currency unit (Rappen), and data scoping. The only notable omission is a pagination/limit/result-size behavior for what is likely a list operation, keeping it from a perfect 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden and largely delivers: it supplies the status enum values (draft | allocated | reversed), defines itemId as 'vouchers that touch that item', and explains that fromDate/toDate define an effectiveDate range. workspaceId is only implied via the tenant-scope note rather than explicitly tied to the parameter name, a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource, 'List landed-cost vouchers', and enriches it with concrete filter semantics (status values, itemId, effectiveDate range). The listing intent is unmistakable and clearly separable from siblings like landed_cost_get (single voucher), landed_cost_voucher_create, and landed_cost_reverse.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when this tool applies: listing vouchers with status/item/date filters, and it notes the tenant-scoping constraint ('A foreign workspace sees only its own vouchers'). It does not explicitly name alternatives or exclusion conditions, but the filter and output detail makes the use case unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
landed_cost_reverseA
Reverse an ALLOCATED voucher, the ONE correction for a confirmed allocation. Nothing about the original voucher, its movements or its journal is edited or deleted. What is written is the compensation: a value-only J02 landed_cost movement per target with the NEGATED cost against the same receipt layer (so J03 nets the layer landed cost back to zero, inventory value flat), and a reversing A02 entry that is the exact mirror of the confirm entry (so the GL nets flat). The voucher becomes reversed and is permanently linked to the reverse journal and reverse movements. reason is required. Idempotent under idempotencyKey. period_locked when the effective period is sealed. CONSEQUENCE: Posts the compensating reversal of a confirmed landed-cost allocation, without editing the original.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| voucherId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it delivers extensively: it explains that the original is untouched, what compensating entries are written (J02, A02), the netting effects, permanent linkage, idempotency under idempotencyKey, and the period_locked condition. This is unusually transparent for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but dense with necessary accounting detail, and it is front-loaded with the core purpose before consequences. Minor redundancy exists between 'Nothing about the original... is edited or deleted' and the closing 'without editing the original,' so it is not maximally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-stakes financial reverse operation with no annotations and no output schema, the description explains the mechanism, resulting state, idempotency, and a failure condition (period_locked). It lacks a fuller list of error scenarios or response shape, but those are not strictly required here given the detail already provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage across four required parameters. The description adds meaning for reason ('reason is required') and idempotencyKey ('idempotent under idempotencyKey'), while voucherId is made clear by the opening sentence. workspaceId is left to convention, and not every parameter gets explicit treatment, but the names and surrounding context make the gaps manageable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Reverse') and resource ('an ALLOCATED voucher') and frames it as 'the ONE correction for a confirmed allocation,' which clearly distinguishes it from sibling tools like landed_cost_allocate_confirm and landed_cost_voucher_create. An agent can tell what this tool does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides strong usage context: this is the correction path after allocation is confirmed, and it does not edit or delete the original voucher. It does not explicitly name alternative tools or state when not to use it, but the state constraint ('ALLOCATED', 'confirmed allocation') plus 'ONE correction' routes an agent effectively.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
landed_cost_voucher_createA
AGENT-FIRST. Open a DRAFT landed-cost voucher that collects freight, duty, insurance, handling or brokerage against one or more POSTED goods-receipt lines, so the extra cost can be capitalised onto the inventory that incurred it (OR 960 acquisition cost). costLines is a list of { componentType (freight | duty | insurance | handling | brokerage | other), amountMinor (> 0 Rappen), description?, vendorId?, taxCodeId? }. targetGrLineIds are the recognised goods_receipt_doc_line ids (each must be on a posted, non-reversed receipt and not already allocated under a live voucher, else already_allocated). inventoryAccountId and clearingAccountId are the two GL accounts a later confirm posts between (Dr inventory control, Cr landed-cost clearing / accrued costs). varianceAccountId (optional) receives the NON-capitalizable share when a receipt has been partly issued or the item is standard-cost (J03 carries only the on-hand share onto inventory); it is required at confirm only if such a remainder actually arises. variancePolicy is expense_excess (default: route the remainder to varianceAccountId) or strict (refuse a confirm that would leave a remainder). allocationMethod defaults to by_value. Nothing physical happens yet: no stock movement, no journal. Idempotent under idempotencyKey.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| fxRate | No | ||
| currency | No | ||
| costLines | Yes | ||
| isEstimated | No | ||
| workspaceId | Yes | ||
| effectiveDate | No | ||
| sourceBillIds | No | ||
| idempotencyKey | Yes | ||
| variancePolicy | No | ||
| targetGrLineIds | Yes | ||
| allocationMethod | No | ||
| clearingAccountId | Yes | ||
| varianceAccountId | No | ||
| inventoryAccountId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and succeeds admirably. It discloses side-effect boundaries ('no stock movement, no journal'), idempotency under idempotencyKey, the already_allocated failure condition, conditional varianceAccountId behavior, and defaults for variancePolicy and allocationMethod.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: purpose, parameter semantics, constraints, defaults, side effects, and failure modes. It is front-loaded with the primary action and does not contain filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 15-parameter accounting tool with no output schema and no annotations, the description provides what an agent needs to call it correctly: preconditions, side-effect boundaries, defaults, error conditions, and the purpose of the involved accounts. No critical operational detail is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does so strongly for core parameters: costLines item shape, targetGrLineIds validity rules, account roles, variancePolicy options, and idempotencyKey. Some optional params like currency, fxRate, effectiveDate, sourceBillIds, and isEstimated get no semantic explanation, leaving a partial gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Open a DRAFT landed-cost voucher' that collects freight, duty, insurance, handling or brokerage against posted goods-receipt lines. It also clearly distinguishes this from later workflow steps by noting 'Nothing physical happens yet' and that a later confirm posts the journal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes clear context: this is the draft-creation step, it is idempotent, and it has no physical effect. It also states preconditions for target GR lines and the variance policy. However, it does not explicitly name sibling alternatives like landed_cost_allocate_confirm or landed_cost_allocate_preview as the when-to-use alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ledger_qaARead-only
Answer a question from the books without opening a report: Umsatz (turnover), offene Posten (open debtors) or MWST (VAT). Reads the existing A08/A16/A07 models and never mutates; a VAT question needs periodStart and periodEnd.
| Name | Required | Description | Default |
|---|---|---|---|
| question | Yes | ||
| periodEnd | No | ||
| periodStart | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes that the tool does not mutate, so the description's 'never mutates' adds little. It does add the behavioral detail that it reads existing A08/A16/A07 models and that VAT questions require period parameters, but it does not disclose the answer/response format or other limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with no filler. The first sentence front-loads purpose and scope, and the second packs behavioral context and a parameter condition efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for selecting the tool but not fully sufficient for reliable invocation: there is no output schema, no parameter descriptions, and the response format, valid question phrasing, and workspaceId semantics are not covered. It gives a good high-level picture but leaves several operational details to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It usefully indicates that the question can be about turnover, open debtors, or VAT, and that periodStart/periodEnd are required for VAT questions. However, it leaves workspaceId, the expected question format, and the exact period value format unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Answer a question from the books') and enumerates the supported subjects: Umsatz, offene Posten, and MWST. The phrase 'without opening a report' differentiates it from report/export siblings even though no sibling is named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when to use the tool: when a direct answer from the books is needed rather than a report. It also provides a conditional usage rule: a VAT question needs periodStart and periodEnd. It does not name specific alternative tools or explicit when-not-to-use scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_accountsARead-only
List accounts, optionally filtered by search text. savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| search | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| includeArchived | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true, and the description adds non-obvious behavior: saved-view filters are merged underneath explicitly named filters. This explains a subtle interaction beyond what the schema reveals. It doesn't address pagination or defaults for includeArchived, but the safety profile is covered by annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences deliver the core action up front and reserve the second sentence for the nuanced saved-view behavior. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main listing behavior and saved-view interaction but omits details about includeArchived semantics and pagination, and there is no output schema. For a four-parameter list tool, a bit more detail would make it complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning for 'search' (search text) and 'savedViewId' (saved view merge behavior), but doesn't describe 'includeArchived' or 'workspaceId'. With half the parameters unaddressed, it's only partially compensatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List accounts') and indicates optional search filtering, making the tool's purpose immediately understandable. It doesn't explicitly compare against sibling list/search tools, but the resource and action are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: it lists accounts and explains how savedViewId interacts with explicit filters. It doesn't state when not to use it or mention alternatives, but the context is sufficient for basic selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_agent_sessionsARead-only
List the agent sessions (Gespräche) of a workspace, newest first, each with its call, write and proposal counts. The trace records at the dispatch seam, so every session an agent seat ever opened is here; from/to filter on last activity, openOnly keeps running sessions, entityRef keeps the sessions whose trace created or approved that one object (the id a verb answered with), which is how a detail view links into the conversation it came from.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| openOnly | No | ||
| entityRef | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true, already covering the safety profile. The description adds valuable behavioral context: the dispatch-seam trace ensures full history, and entityRef semantics ('the id a verb answered with') explain the linking behavior. It does not contradict annotations and goes beyond the schema's bare parameter names, though it omits pagination or limit details, which are not critical for a list tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences pack in the core purpose plus all parameter semantics. The main action is front-loaded, and every clause adds information (filter semantics, count fields, trace explanation). No filler or redundant content. The compact structure is appropriate for the density of information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with five parameters, no output schema, and a rich interaction environment, the description covers the purpose, filters, ordering, counts, and linking behavior. It omits return-format specifics beyond counts (e.g., whether fields like lastActivity are present), but the core invocation needs are fully addressed. Slight deduction for not clarifying date format for from/to, though that is inferable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden of explaining parameters. It explains every parameter: from/to filter on last activity, openOnly keeps running sessions, entityRef matches the object created/approved by a trace, and workspaceId is the scope. This exceeds baseline compensation expectations and adds meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb ('List'), resource ('agent sessions'), scope ('of a workspace'), ordering ('newest first'), and what each entry includes ('call, write and proposal counts'). It also clarifies the trace semantics (every session an agent seat ever opened) which distinguishes this list tool from singular session getters like get_agent_session. The purpose is unambiguous and differentiates from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context on when to use the tool and what each filter does, including concrete use cases like 'how a detail view links into the conversation it came from'. It does not explicitly name alternatives (e.g., get_agent_session) or state exclusions, but the list vs. singular distinction is implicit. This meets the 'clear context, no exclusions' level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_automation_rulesARead-only
List this workspace automation rules, newest first, archived ones excluded unless asked for. Also returns the CATALOGUE a rule can be built from: every registered trigger event, every write verb that is a legal action, and every condition operator, all read off the live registries so a picker cannot offer something the engine would refuse. savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| event | No | ||
| enabled | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| includeArchived | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses ordering (newest first), default archive exclusion, the live catalogue source, and the saved view filter merging behavior. These traits go well beyond the readOnlyHint annotation, providing rich behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the main purpose, then catalogue details, then saved view nuance. Every sentence adds concrete value with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main outputs (rules and catalogue), ordering, filtering, and saved view behavior. It omits explicit semantics for event and enabled, though their names are self-explanatory. Given the absence of output schema, this is fairly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explains savedViewId merging and implies includeArchived defaults, but leaves event and enabled parameters undefined. With 0% schema coverage, it partially compensates but does not cover all parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as listing workspace automation rules with ordering and default filtering, and adds the catalogue feature. It does not explicitly differentiate from sibling tools like list_automation_runs or get_automation_rule, so it stops short of full sibling distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives, and no exclusions or alternative tool names are mentioned. The description only covers parameter behavior, not tool selection context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_automation_runsARead-only
The run log, newest first: what fired, when, why, what it actually sent, and what came back. A failed run carries the target verb's own rejection code verbatim, so the reason reads the same way it would on that verb's own screen.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| ruleId | No | ||
| status | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description doesn't need to restate that. It adds useful behavioral details: the order (newest first), the fields included, and that failed runs carry the target verb's rejection code verbatim. This is beyond the structured data and helps the agent understand output semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences long, with no fluff. The key information is front-loaded: what the tool returns and its ordering. It is efficient and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a list tool, but lacks details about the return format (e.g., how many runs per page, whether it's an array), how the parameters affect results, and when to prefer this over get_automation_run. Since there is no output schema, more context would help the agent know what to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 4 parameters (workspaceId required, limit, ruleId, status) but the description gives no explanation of their meaning or usage. With schema description coverage at 0%, the description should compensate, but it doesn't mention any parameters at all. The parameter names are somewhat self-explanatory, but for limit and status, the expected values or formats are not hinted at.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states what the tool does: it lists the run log for automations, newest first, and details what each entry contains (what fired, when, why, what was sent, and what came back). This distinguishes it from the sibling get_automation_run, which presumably retrieves a single run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is for viewing a log, but it does not explicitly state when to use this list tool versus get_automation_run. It also doesn't mention filtering options or when to use the ruleId/status parameters. The usage context is clear enough for a log listing, but lacks explicit alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_backupsARead-only
List this workspace`s backups and exports, newest first, with size, timestamp, actor and status. Optional kind (backup|export) or status (complete|failed) narrows the list. savedViewId applies a saved preset (G00): its stored filters are merged underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| status | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds ordering and savedViewId filter-merging behavior. It does not disclose potential pagination, rate limits, or permission requirements beyond the read-only hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences: first states purpose and output fields, second explains filters. No fluff, front-loaded with the core operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with 4 params and no output schema, the description covers the essential inputs and output fields. It lacks explicit pagination/limit info, but given the tool's simplicity, it is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain all parameters. It explains kind with allowed values (backup|export), status (complete|failed), and savedViewId merging semantics. workspaceId is obvious from context, so all four parameters are effectively covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it lists backups and exports with specific fields (size, timestamp, actor, status) and ordering (newest first). It distinguishes from create/delete/restore siblings implicitly, but does not explicitly name alternatives like list_restorable_backups.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides filter usage for kind and status, but gives no explicit guidance on when to use this tool versus related siblings (e.g., list_restorable_backups). No exclusions or alternative routing are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_bank_accountsARead-only
List this workspace's Bankkonten with their IBAN, currency, verknüpftes Konto, QR-IBAN flag and opening balance. Archived accounts are excluded unless includeArchived is true. savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| savedViewId | No | ||
| workspaceId | Yes | ||
| includeArchived | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description aligns with a read-only listing operation. It adds meaningful behavioral context: archived accounts are excluded by default, includeArchived overrides that, and savedViewId merges stored filters underneath explicit filters. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all information-dense and front-loaded. The core listing purpose and returned fields come first, then filtering behavior and saved-view semantics. No filler or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list with 3 parameters, the description covers the purpose, returned fields, archive filtering, and saved-view merging. There is no output schema, so the listed returned fields are helpful. It could optionally note that workspaceId is the required scope, but that's implied by 'this workspace's' and the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden. It explains includeArchived semantics explicitly, and gives a useful preview of savedViewId behavior including the G00 reference and filter merging. It doesn't describe workspaceId, but that parameter is self-evident from the tool name and required field. This is strong compensation for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly specifies the verb (list), the resource (this workspace's Bankkonten), and the key fields returned (IBAN, currency, verknüpftes Konto, QR-IBAN flag, opening balance). It distinguishes the scope to 'this workspace' but doesn't explicitly contrast it against sibling list_bank_statements or the get_bank_account tool, so it misses a bit of differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states that archived accounts are excluded unless includeArchived is true, providing a concrete condition for a parameter. It also explains how savedViewId interacts with filters. However, it doesn't explicitly say when to use this tool over get_bank_account or list_bank_statements, though the context is fairly clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_bank_statementsARead-only
The imported camt statements of the workspace, newest period first, each with its open-line count (lines still unmatched or partial) and the D64 reconciled indicator. This is the door to a statement: take its statementId into list_reconciliation and suggest_matches. bankAccountId narrows to one Bankkonto. Writes nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| bankAccountId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description's 'Writes nothing' adds emphasis but not contradiction. It adds context beyond annotations by explaining open-line count meaning and the D64 reconciled indicator, and clarifies the ordering (newest first). However, it doesn't describe pagination or output format, which is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but concise, packing scope, ordering, derived fields, and downstream usage into two sentences. It front-loads the main purpose and immediately provides actionable routing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with two params and no output schema, the description covers the key information an agent needs to invoke it correctly and follow up. It could mention pagination or result limits, but the provided detail about open-line counts and D64 is sufficient for the core task. The absence of output schema is mitigated by the description's clarity on what each statement includes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description is the only source of parameter meaning. It explains that bankAccountId narrows to one Bankkonto, which adds context beyond schema names. It does not detail the workspaceId format or whether it's required, but the schema marks it required, and the description implicitly relies on that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists imported camt statements for the workspace, newest first, with open-line count and D64 indicator. It distinguishes itself by explicitly mentioning the statementId as the door to list_reconciliation and suggest_matches, and highlights the bankAccountId narrowing option, which differentiates it from generic list tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states this is the entry point to statements and directs the agent to take statementId into the next steps (list_reconciliation and suggest_matches). It also clarifies the optional parameter for narrowing by bank account and confirms it writes nothing, which guides when to use it versus other tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_capturesBRead-only
The Belegeingang queue: captures with their live proposed fields. The default view excludes discarded captures; pass status to filter to needs_review, committed or discarded. from/to filter on capture date; savedViewId applies a G00 saved view.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| status | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations include readOnlyHint=true, which is consistent with the description (listing captures). Adds useful details like the default exclusion of discarded and date filtering, but does not explain return format or pagination, though output schema absence leaves some vagueness.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the main purpose. Each sentence addresses a key aspect. Slightly dense but no superfluous wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity and the annotations covering read-only, the main gaps are parameter semantics and lack of output/return details. The description covers default behavior, status filtering, and date range, but more explicit enum values for status and clarification of workspaceId would help.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters. It mentions 'status' (with possible values), 'from/to' as date filters, and 'savedViewId' as a saved view, but does not clarify 'workspaceId' or the string formats for dates/status. This is partial compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool lists captures from the Belegeingang queue with proposed fields, and specifies default view behavior. While it doesn't distinguish from all siblings (e.g., get_capture, capture_commit), it conveys the core purpose effectively.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides specific guidance on filters: default excludes discarded, status and from/to filters, and savedViewId. Does not explicitly mention when to prefer alternatives, but the sibling list shows capture_document and capture_commit, and the description implies this is for listing rather than manipulating.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_conceptsARead-only
List the Begriffe: TILL's authored explanations of its Swiss accounting vocabulary (Saldosteuersatz, Vorsteuer, Steuerperiode, ...). Filter by area or a per-token query; returns key, localized term and area, never the bodies. Workspace-free: the corpus is identical in every workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| area | No | ||
| query | No | ||
| locale | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds meaningful behavioral detail beyond that: it returns only key, localized term, and area, never the bodies, and is workspace-free. This helps agents understand the lightweight, non-mutating nature of the call and its cross-workspace consistency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core purpose, and every sentence contributes: what the tool lists, how to filter, what is returned, and the workspace-independence guarantee. No filler or redundant restatement of the tool name or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with three optional parameters and no output schema, the description covers the main return shape and filtering behavior well. The only gap is the undocumented locale parameter and slightly vague 'per-token query' wording, but overall this is nearly complete for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain 'area' and 'per-token query' filters, but it does not explain the locale parameter at all, nor does it define the exact semantics of 'per-token query.' This is partial compensation for the undocumented schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('List') on a well-defined resource ('TILL's authored explanations of its Swiss accounting vocabulary') and clarifies that it returns metadata only ('key, localized term and area, never the bodies'). This clearly distinguishes it from the sibling get_concept without needing to open the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete usage context: filter by area or a per-token query, and notes that the corpus is workspace-free, which tells agents when scope does not matter. It does not explicitly name alternatives or say when not to use this tool, but the 'never the bodies' phrasing implies the boundary against body-retrieval tools like get_concept.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_contactsARead-only
List contacts, filtered by query or party role. savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| partyRole | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| includeArchived | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already signals a safe read operation, so the bar for additional disclosure is lower. The description adds useful behavioral nuance about savedViewId: stored filters are merged underneath explicitly named filters, which goes beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with no filler. The primary action and filter options are front-loaded, and the savedViewId nuance earns its place because it is non-obvious behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool, the description covers the core invocation intent and the most nuanced parameter behavior. It omits return-shape and pagination details, but the readOnly annotation and the straightforward 'list contacts' intent keep it reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaning for query, partyRole, and savedViewId, including the merge behavior. With 0% schema description coverage it partially compensates, but workspaceId and includeArchived remain undocumented in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'List contacts', with filtering by query or party role. It is clear enough to be distinguishable from sibling contact tools like contacts_timeline or contacts_log_activity, though it does not explicitly call out those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when to use the tool: when listing contacts with optional query, party-role, or saved-view filtering. It does not state exclusions or name alternative tools, but the context is not ambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_cost_centersARead-only
List cost centres. savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| savedViewId | No | ||
| workspaceId | Yes | ||
| includeArchived | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true, so the description adds value by disclosing the saved view filter merging behavior, which is non-obvious. It does not describe return format or pagination, but for a read-only list tool this is sufficient added context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words. The core action is front-loaded, and the saved view behavior is explained in a single follow-up sentence. Perfectly sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with readOnlyHint annotation, the description covers the primary purpose and the key tricky parameter. It omits default behavior for includeArchived and pagination, but these are not critical for invoking the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaningful semantics for savedViewId (the most ambiguous parameter) by explaining how its filters merge. workspaceId and includeArchived are self-explanatory from their names and typical usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'List cost centres' with a clear verb and resource, and the saved view behavior is a helpful extra. It is easily distinguished from sibling tools like create_cost_center, archive_cost_center, and delete_cost_center.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains how savedViewId behaves but gives no explicit guidance on when to use this tool versus alternatives. There are no mentions of exclusions or when to prefer other list tools, so usage context is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_dispatchesARead-only
Das Versand-Protokoll: the cross-document send log, newest first, one row per recipient per send attempt (send_invoice, send_dunning_run, quotes_send), each carrying kind, document/run link, contact, recipient, channel (smtp|cloud_relay|artifact_only), locale, the resolved subject and body that actually left, outcome (sent|degraded|failed|artifact_created) and a degrade reason. Filter by kind, contact, outcome or date range; savedViewId applies a saved view (G00). Also returns texts, the saved Textbausteine slots, so the editor reads its state in the same call. Append-only: no verb can edit or delete a row.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| outcome | No | ||
| contactId | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| documentKind | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the readOnlyHint annotation by stating 'Append-only: no verb can edit or delete a row', reinforcing the read-only nature and adding a critical behavioral constraint. It also discloses ordering ('newest first'), the fact that it returns the resolved subject/body that actually left, and that it also returns the saved Textbausteine slots for editor state. These details are not in the annotation and significantly enrich behavioral understanding.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-rich, front-loading the core purpose and then enumerating return fields and filtering options. It is a single paragraph but well-organized, covering all essential aspects without excessive verbosity. It could be improved by breaking into sections or bullets, but it is not overly long relative to the information conveyed. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (7 params, no output schema), the description is quite complete. It explains what the tool returns (rows with fields, including the texts for editor state), the append-only nature, ordering, and filtering options. The only notable omission is pagination or limit behavior (e.g., whether all rows are returned or truncated), which could be important for an agent expecting large result sets. Overall, it covers the essential context for correct invocation and interpretation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 7 parameters with 0% description coverage in the schema itself, so the description must compensate. It mentions filtering by 'kind, contact, outcome or date range' and the savedViewId, which maps to documentKind, contactId, outcome, to/from, and savedViewId respectively. However, it does not explicitly map each parameter name to its semantic meaning (e.g., 'to' and 'from' as date bounds). It provides general guidance but leaves some ambiguity about exact parameter roles, so it adds value but does not fully compensate for the schema's lack of descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: it lists dispatches from a cross-document send log, with explicit detail about what each row contains (kind, document/run link, contact, recipient, channel, locale, resolved subject/body, outcome, degrade reason) and which operations generate entries (send_invoice, send_dunning_run, quotes_send). The verb and resource are specific, and the description distinguishes it from a generic list by noting the append-only nature and the inclusion of saved text slots. It is unambiguous and easily distinguishable from siblings even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on when to use this tool: to view the send log, filter by kind, contact, outcome, or date range, and apply a saved view. It also implies it's for inspecting dispatch history. However, it does not explicitly state when not to use it or name alternative tools (e.g., list_dunning_runs for dunning-specific logs). The guidance is sufficient for an agent to infer appropriate use, but lacks explicit exclusions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_documentsARead-only
List documents, filtered by type, status, contact, credited invoice (A13: the credit notes of invoice X), or date range. savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here, so an explicit one always wins.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| type | No | ||
| status | No | ||
| contactId | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| creditedDocumentId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the agent knows this is a safe read operation. The description adds value by explaining the savedViewId filter-merging behavior and the credited invoice filter, which are non-obvious. However, it does not disclose pagination, ordering, or return format, leaving some behavioral gaps. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences that front-load the core purpose ('List documents') and then efficiently explain the savedViewId merging behavior. Every sentence carries essential information, with no fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (8 parameters, no output schema), the description covers the main filtering options and the savedViewId behavior, but it omits return structure, pagination, and filter combination semantics. It is adequate for basic usage but incomplete for an agent that needs to know exactly what results look like or how to combine filters correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain the purpose of most parameters: type, status, contact (contactId), credited invoice (creditedDocumentId), and date range (from/to), plus savedViewId. However, it does not clarify value formats (e.g., date format), whether filters are ANDed, or the role of workspaceId beyond being required. It adds meaning beyond the schema but leaves important details unspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists documents with multiple specific filters (type, status, contact, credited invoice, date range) and explains the savedViewId behavior. It is distinct from sibling tools like get_document (single retrieval) and list_document_templates (templates), so an agent can identify its purpose without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus alternatives (e.g., get_document for a single document, search_global for general search). It implies usage for listing documents but provides no exclusions or conditions for choosing this over other list tools. The savedViewId merging nuance is helpful but does not address tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_document_templatesARead-only
List document templates, newest defaults first per kind: id, kind, name, default/archived flags, footer locales, language mode, and the linked logo file id. documentKind narrows to one kind; includeArchived includes archived rows; savedViewId applies a saved view (G00).
| Name | Required | Description | Default |
|---|---|---|---|
| savedViewId | No | ||
| workspaceId | Yes | ||
| documentKind | No | ||
| includeArchived | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, so the description doesn't need to state read-only. It adds behavioral context by specifying the sort order, the set of returned fields, and how parameters alter the result set. This goes beyond the annotation and provides transparency about what the call returns and how it behaves, warranting a 4.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, first states the action and output, second explains parameters. No filler. Front-loaded with the core purpose. This is concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list operation with no output schema, the description provides the fields returned, sorting, and parameter semantics. It doesn't mention pagination or limits, which might be a minor gap, but given the simplicity of the operation, it's adequately complete. A 4 reflects that it covers the essentials but could include a note on result size or pagination.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must explain parameters. It does: documentKind narrows to one kind, includeArchived includes archived rows, savedViewId applies a saved view (G00). It also implies workspaceId is the workspace scope. This fully compensates for the schema gap, so a 5 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists document templates, specifies the fields returned (id, kind, name, flags, locales, language mode, logo file id), and the default sorting order (newest defaults first per kind). It distinguishes from get_document_template by its list nature and from list_documents by focusing on templates. This is a specific verb+resource with clear scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains how each parameter filters results (documentKind, includeArchived, savedViewId), which is useful. However, it does not explicitly state when to use this tool versus alternatives like get_document_template or list_documents, nor does it mention any prerequisites. The usage is implied but not fully explicit, so it scores a 3.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_drafted_actionsARead-only
List the drafted agent actions (Vorschläge) of a workspace: pending by default (oldest first), or executed/rejected via status. Each row carries its verb, parsed payload, dial capability, proposer and resolution, so the approval queue renders from one read on every face.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description is consistent with that. It adds valuable behavioral details beyond the annotation by specifying default ordering (oldest first) and the ability to filter by status. It also describes the row content, which helps set expectations for the response. Without mentioning limits or pagination, it still adds useful context beyond the structured annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no wasted words. The core purpose and default behavior are front-loaded, followed by the output row details and the intended use case. It is succinct and well-structured, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with one required and one optional parameter, the description is fairly complete. It explains the output fields (verb, payload, dial capability, proposer, resolution) and the intended use. It does not mention pagination or API limits, but given the simplicity of the tool and the readOnly annotation, this is a minor gap. Overall, it provides enough information for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and parameters have no descriptions. The description compensates by explaining the semantic role of both parameters: workspaceId scopes the list to a workspace, and status overrides the default pending filter with executed/rejected. It does not list exact enum values, but the behavior is clear. This is a good compensation given zero schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it lists drafted agent actions for a workspace, and distinguishes itself from sibling mutation tools like approve_drafted_action and reject_drafted_action by focusing on the read operation. It also explains the default status (pending) and ordering, leaving no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives context on when to use the tool (for rendering the approval queue) and explains how the status parameter filters results (executed/rejected vs. default pending). It does not explicitly mention when not to use it or compare to alternatives, but the sibling tools clearly imply this is for listing before acting, so usage is reasonably clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_dunning_runsARead-only
The Mahnlauf history, newest first: each run's date, status (proposed/issued/sent), item and debtor counts, highest level, and whether a fee was booked or skipped by a period lock. Filter by status or date range. savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here. Reads only.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| status | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation is reinforced and supplemented with useful behavioral detail: newest-first ordering, the status vocabulary, counts, highest level, period-lock fee behavior, and saved-view filter merging. Pagination or result-limit behavior is not mentioned, but for a read-only list this is strong coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences deliver the core purpose, output contents, ordering, filters, saved-view behavior, and read-only status. Information is front-loaded and every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers what the return data represents, the ordering, available filters, saved-view behavior, and safety profile. Minor gaps such as pagination or exact field names prevent a perfect score, but the tool is fully usable from this description alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates well: it explains date-range filtering for from/to, status values, and the savedViewId merge semantics. It does not explicitly describe workspaceId or date formats, but those are substantially self-explanatory or implied by the required schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: listing the Mahnlauf/dunning-run history. It also specifies ordering ('newest first') and the exact data returned per run, which clearly distinguishes it from singular tools like get_dunning_run or actions like propose_dunning_run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly communicates when this tool is appropriate: listing dunning-run history and filtering by status or date range. It does not explicitly name sibling alternatives or exclusion conditions, but the 'history' framing and filter options give clear contextual usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_exchange_ratesARead-only
List recorded foreign-currency rates, newest validity date first, optionally filtered by currency pair and date range.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| workspaceId | Yes | ||
| baseCurrency | No | ||
| quoteCurrency | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and the description's 'List recorded' wording is consistent with a read operation, so there is no contradiction. The description adds useful behavioral context by specifying the newest-validity-date-first ordering, but it does not disclose pagination, permissions, or return-value behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one sentence with no filler: it leads with the core action, resource, and ordering, then appends the optional filtering scope. It is appropriately sized and immediately scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter tool with no output schema and zero schema coverage, the description leaves critical invocation details undocumented: which parameters represent the currency pair and date range, required workspaceId semantics, and expected input formats. The clear purpose alone is not enough for an agent to reliably construct correct calls.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only vaguely mentions 'currency pair and date range' without mapping those concepts to the five parameters. It is unclear whether from/to mean dates or currencies, or how baseCurrency and quoteCurrency relate to the pair. This ambiguity gives the agent insufficient guidance for populating required and optional fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource—'List recorded foreign-currency rates'—with explicit ordering ('newest validity date first') and optional filtering. These semantics distinguish it from sibling tools like get_exchange_rate, record_exchange_rate, and import_exchange_rates, even though no siblings are named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for retrieving a collection of existing rates, especially when newest-first ordering or filters are desired. However, it never explicitly names alternative tools or states when not to use this one. An agent must infer the distinction between listing rates versus getting, recording, or importing them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_feedbackCRead-only
List the feedback reports written on this computer. Every one reads "prepared": TILL hands a report to a mail client and cannot observe whether it was sent.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the read-only nature is covered. The description adds the nuance that reports show 'prepared' status because TILL cannot confirm if sent, which is valuable behavioral context beyond the annotation. However, it's cryptic (unclear what TILL means) and doesn't explain other behavioral aspects like sorting or filtering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, so it's concise. However, the second sentence is awkwardly phrased and likely confusing ('TILL' appears to be a typo or undefined acronym). The structure doesn't front-load the most important information about parameters or output. It's short but not well-organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With only one parameter and no output schema, the description still leaves major gaps: it doesn't explain what a feedback report contains, what the return format looks like, or whether there's pagination. The 'prepared' status semantics are partially covered but incompletely. For a simple list tool, more context about the output and the meaning of statuses is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, workspaceId, has 0% schema description coverage, and the description provides no explanation of what workspaceId represents, its format, or its meaning in this context. With zero coverage, the description should compensate, but it doesn't. This is a critical gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it lists feedback reports and scopes them to 'this computer'. The verb 'list' plus the resource 'feedback reports' is specific. It doesn't explicitly differentiate from sibling tools like prepare_feedback or preview_feedback, but the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. It doesn't mention related tools like preview_feedback or prepare_feedback, nor does it state conditions for selecting this over them. The only context is the status explanation, which is behavioral, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_field_defsARead-only
List the custom fields defined on an entity kind, in display order. Archived fields and unconfirmed drafts are excluded unless asked for.
| Name | Required | Description | Default |
|---|---|---|---|
| entityKind | Yes | ||
| workspaceId | Yes | ||
| includeDrafts | No | ||
| includeArchived | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint=true is provided, and the description correctly indicates a read operation (list). The description adds value by disclosing the default filtering behavior (excluding archived and drafts), which is not in structured data. However, it does not elaborate on pagination, sorting guarantees beyond display order, or potential side effects of including drafts/archived, so a mid-range score is apt.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence plus a second short sentence, extremely concise and front-loaded. Every clause conveys essential information: purpose, ordering, and default exclusions. There is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (4 params, no output schema), the description covers the key behavioral aspects: listing definitions, ordering, and default exclusions. The absence of an output schema means the description is not expected to detail return values. A minor gap is not stating whether drafts and archived fields are included when explicitly requested, but the description's negation implies that. Overall, comprehensive enough for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% (the input schema has no descriptions for parameters), so the description must compensate. The description explains the core purpose of workspaceId and entityKind through the phrase 'entity kind' and context. It also explains the boolean flags' role via 'Archived fields and unconfirmed drafts are excluded unless asked for,' directly mapping to includeArchived and includeDrafts. This adds significant meaning beyond the bare parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear and specific: 'List the custom fields defined on an entity kind, in display order.' It specifies the verb (list), resource (custom fields), parent resource (entity kind), and output ordering. This distinguishes it from siblings like list_field_values, which list values rather than definitions, and define_field, which creates fields.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear context: operations are on an entity kind, and the description notes exclusions by default ('Archived fields and unconfirmed drafts are excluded unless asked for'). This implies when to use the includeDrafts and includeArchived parameters. It does not explicitly mention alternatives, but the description's focused scope makes usage apparent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_field_valuesARead-only
Read every custom field value stored on one record, archived fields included and labelled as archived. This is how a retired field's data stays reachable after it stops rendering as an input.
| Name | Required | Description | Default |
|---|---|---|---|
| entityId | Yes | ||
| entityKind | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already signals safety, and the description confirms the read-only nature while adding a behavioral detail not in the annotations: archived field values are returned and marked as archived. This gives agents useful expectations beyond the structured annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler; the read action and scope are front-loaded, and the archived-field behavior earns the second sentence. Every phrase adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only listing tool with no output schema, the description covers what is returned, including the archived-field caveat. The main gap is the lack of parameter-level explanation, but the core call context is sufficiently clear.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the undocumented workspaceId, entityKind, and entityId parameters, but it only indirectly references 'one record.' It does not explain accepted entityKind values, how the three parameters relate, or any value formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with the specific verb and resource: 'Read every custom field value stored on one record,' and adds the distinguishing scope that archived fields are included and labelled. This clearly separates it from write tools like set_field_value and definition tools like list_field_defs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The second sentence establishes the main use case for this tool: retrieving data from retired fields that no longer render as inputs. It gives clear context for when to call it, though it does not explicitly name alternatives or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_itemsCRead-only
List items, filtered by query. savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| includeArchived | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the read-only safety is covered. The description adds a useful behavioral nuance: savedViewId merges its stored filters beneath explicitly provided filters. However, it omits other behavioral aspects such as handling of includeArchived or result set ordering, leaving the description to partially carry the burden despite the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with the main operation first, followed by the savedViewId nuance. The cryptic '(G00)' reference adds slight noise, but the description is efficient and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lacks essential context for a tool with 4 parameters incl. a required one and no output schema. It doesn't explain what kinds of 'items' are listed, how includeArchived affects results, the need for workspaceId, or any pagination/return format. This is incomplete for an agent to call the tool correctly without external knowledge.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It offers meaningful semantics for savedViewId (applies a saved view and merges filters) and implies query's filtering role, but workspaceId (required) and includeArchived are left undefined. This is insufficient for a 4-parameter schema with no other parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear operation ('List items') and a filtering mechanism ('filtered by query'). The resource 'items' is distinct from sibling list_* tools (list_accounts, list_contacts), though it doesn't explicitly differentiate itself. The savedViewId detail adds precision without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to choose this over alternative list tools, nor any exclusions or prerequisites. The savedViewId explanation is a behavioral detail, not a usage guideline. With numerous sibling list_* tools, the lack of selection criteria is a notable gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_journalARead-only
List journal entries, filtered by date range, account, source, or status. Each entry reports the currency its total is in (null when the entry has no lines), plus the base-currency total, rate and base currency for a foreign-currency entry, all derived from its posted rows. savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| source | No | ||
| status | No | ||
| account | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already signals a safe read operation, and the description adds meaningful behavioral context: entries report currency (null when no lines), base-currency total/rate, and base currency derived from posted rows, plus the savedViewId merge behavior. It does not mention pagination, ordering, or limits, but the additional detail goes well beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loaded with the core purpose, followed by output details and the saved view nuance. Every sentence earns its place with no redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 7 parameters and no output schema, the description covers the return-value essentials (currency, total, rate, base currency) and saved view behavior, but omits pagination, ordering, date format, and potential response envelope. This is adequate for a moderately complex read-only list tool, but an agent would still need to infer some conventions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the burden of explaining parameters. It maps 'date range' to to/from, and account, source, status as filters, and explains savedViewId's merge behavior, providing useful meaning. However, it does not clarify value formats, allowed statuses/sources, or workspaceId semantics, leaving gaps for an agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List journal entries') and names the filter dimensions (date range, account, source, status). It is clear, but it does not explicitly differentiate from siblings like get_entry, export_journal, or general_ledger, so it falls short of full sibling distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives such as get_entry for a single entry or export_journal for exporting. The filter list implies some usage contexts, but there are no exclusions, alternatives, or prerequisite conditions, leaving an agent to infer when this tool is preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_membersARead-only
List everyone bound to this workspace, pending invites included and labelled as pending (a pending member holds no capability at all).
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true. The description adds meaningful behavioral context beyond that: pending invites are included and labelled, and a pending member holds no capability. This is extra information not present in the schema or annotations, helping the agent set expectations. No contradictions exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence. It immediately states what the tool does patron, then adds essential nuances (pending invites, capability). Every word adds value with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity—one required parametercars, no nested objects, read-only annotation—the description covers the core behavior and output expectations. It lacks mention of pagination or result format, but those are minor for this straightforward listing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage for the required workspaceId parameter. The description indirectly references it as 'this workspace' but does not explicitly define the parameter or its constraints. For a single obvious parameter, this is adequate but does not fully compensate for the missing schema guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the operation ('List'), the resource ('everyone bound to this workspace'), and important scoping details ('pending invites included and labelled as pending'). It is unambiguous and distinct from sibling list tools like list_contacts or list_bank_statements because it specifies 'everyone bound to this workspace'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly tells the agent when to use this tool (to list all workspace members), but it does not explicitly mention alternatives or explain when not to use it. Given no close sibling tool with a similar purpose, the implicit context is sufficient but not exemplary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_open_itemsARead-only
The OP-Liste: every unpaid or partly-paid invoice with its open amount, due date, days overdue and aging bucket, plus every parked customer payment. A parked payment carries its direction: incoming is a Guthaben the customer holds and reads NEGATIVE, outgoing is a refund not yet matched and reads POSITIVE. Reports bucket subtotals in the invoice currency (bucketTotals) and in base currency (baseBucketTotals, the only one that is an amount when several currencies are open), the grand total both ways, and whether the base total reconciles to account 1100 Debitoren as of the same date. Pass asOf to see the receivables exactly as they stood on a past cut-off (payments after it are ignored). Filter by customerId or currency to narrow the table without changing what reconciled means. Reads only.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| currency | No | ||
| customerId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true and the description reinforces 'Reads only.' It goes beyond the annotation by describing sign conventions for parked payments (incoming reads NEGATIVE, outgoing POSITIVE), the effect of asOf on payment inclusion, and how reconciliation relates to account 1100. This is rich behavioral disclosure that helps an agent interpret results correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and front-loaded with the core definition, but it runs long. Several sentences explain nuances (parked payment direction, bucket subtotals vs. base currency totals) that could be tighter, yet each adds real interpretive value. It is structured as a flowing paragraph rather than scannable, but no sentence is pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool without an output schema, the description gives the agent naming conventions, behavior around asOf, currency handling, and reconciliation context. It does not enumerate every return field, but it characterizes the result shape well enough for an agent to invoke it and interpret the response. Minor gaps: no explicit mention of pagination or maximum result size.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the burden of parameter semantics. The description explicitly explains asOf ('pass asOf to see the receivables exactly as they stood on a past cut-off'), customerId and currency as filters, and workspaceId is implied by the resource context. It stops short of providing format or validation details for each parameter, but for a list tool with simple parameters this is solid coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a read-only listing of open receivables: unpaid or partially-paid invoices with aging details, plus parked customer payments. It names specific output elements (due date, days overdue, aging bucket, bucket totals) and the filtering parameters, which distinguishes it from general list tools like list_payments or customer_balance that likely return different data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use it (to see the OP-Liste / open items) and how to narrow results using customerId and currency. It mentions the asOf parameter for historical cutoff. It does not explicitly contrast with alternatives like aging_report or customer_balance, but the scoping language and detail on 'what reconciled means' give practical context for use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_payableARead-only
The bills a payment run could pick up: every posted A17 vendor bill still open, each flagged with whether its vendor has a creditor_bank_profile, whether that IBAN is a QR-IBAN, what reference kind the bill's own vendorReference classifies as (qrr/scor/free_text/none), whether the currency is one create_payment_batch admits (CHF/EUR), and whether the bill already sits in a live (not yet paid) batch. dueBy filters to bills due on or before a date; vendorId to one vendor.
| Name | Required | Description | Default |
|---|---|---|---|
| dueBy | No | ||
| vendorId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes that this is a read-only operation. The description adds meaningful behavioral detail about what gets returned: each bill is flagged for creditor bank profile, QR-IBAN, reference kind, eligible currency, and live batch presence. It does not mention pagination or sorting, but the core read-only behavior is well disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single long run-on sentence that packs a lot of detail into a dense structure. It is front-loaded with the core purpose, but the list of flags and criteria would be much easier to parse if split into sentences or bullets. All content is relevant, but readability suffers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema, the description explains most return semantics: eligibility, per-bill flags, and available filters. Gaps include undocumented workspaceId, unspecified date format, and lack of pagination/sorting details, but these are minor relative to the rich information actually provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains that dueBy filters to bills due on or before a date and vendorId filters to one vendor, but workspaceId—a required parameter—is not described at all, and the expected date format for dueBy is not specified. This is partial compensation over an empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as listing bills a payment run could pick up, with specific eligibility criteria (posted A17 vendor bills still open) and a detailed list of returned flags. This distinguishes it from siblings like list_vendor_bills, list_payment_batches, and preview_payment by emphasizing the payment-run eligibility perspective.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: this tool is for seeing which open bills a payment run could pick up, and it explains the two filters (dueBy, vendorId). However, it does not explicitly name alternatives or state when not to use it, so it falls short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_payment_batchesARead-only
The payment-run history, newest first: every batch with its status, control sum and item count. status filters the lifecycle (draft/generated/paid). savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here (e.g. 'Batches awaiting bank confirmation' as status=generated, 'Paid this quarter' as status=paid with a date range on the client).
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint, the description discloses useful behavior: results are 'newest first', each batch includes its status/control sum/item count, and savedViewId filters are 'merged underneath' explicitly named filters. This clarifies non-obvious merge semantics rather than merely restating the annotation. It does not claim any destructive capability, consistent with readOnlyHint=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences and front-loads the core behavior before going into parameter details. The saved-view examples add length but earn their place by clarifying how filters combine. It is dense but not bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with three parameters and no output schema, the description covers the returned fields, default ordering, status filtering, and saved-view behavior. It does not discuss pagination or explicitly differentiate from payment-batch siblings, but the essential invocation semantics are present. Given the readOnlyHint and simple schema, this is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description compensates by explaining status values ('draft/generated/paid') and the savedViewId merge behavior with concrete examples. The required workspaceId is not described, but its role is evident from the name and required flag. It adds meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a list operation over payment-run history and specifies returned fields ('status, control sum and item count'), so an agent can infer the resource being read. It does not explicitly distinguish itself from sibling tools like get_payment_batch or list_payments, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement about when to choose this tool over related payment-batch tools (get_payment_batch, payment_batch_transmit, mark_batch_paid) or how it differs from list_payments. The filter semantics for status and savedViewId are described, but no when-to-use guidance or exclusion of alternatives is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_paymentsARead-only
List payments, filtered by direction, status, date range, or the document they settled. savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here, so an explicit one always wins.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| status | No | ||
| direction | No | ||
| documentId | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint already declares the operation is safe, and the description adds genuine behavior beyond that: savedViewId merges stored filters underneath explicitly named filters, with explicit filters taking precedence. No mention of pagination or limits, but the merge precedence is the non-obvious behavior that matters most here.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler: the first states the core action and filters; the second explains the non-obvious saved-view precedence. Key information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The definition covers core filter semantics and saved-view behavior, and readOnlyHint covers safety. However, with no output schema, no enums, and zero parameter descriptions, the agent still lacks allowed values/formats, result pagination, and clear sibling routing, so it is not fully complete for a 7-parameter list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description provides the only semantics: direction, status, date range, and the settled document map to direction, status, from/to, and documentId. It also explains savedViewId's merge semantics. It stops short of giving allowed values or date formats, but it compensates for the empty schema meaningfully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with 'List payments' – a specific verb and resource – and enumerates concrete filter dimensions (direction, status, date range, settled document). This separates it from single-payment getters like get_payment and transactional tools like record_payment/reverse_payment without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The filter list and saved-view merge rule imply when the tool is appropriate, but the description never names alternatives or conditions for preferring another sibling such as get_payment or list_payment_batches. Usage context is present by implication, not explicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_payroll_handoffsARead-only
The payroll hand-off history: the union of exports and wage-journal postings, each row typed export or posting, newest first. An export row carries counts, the AHV-inclusion state and the artifact link; a posting row carries the entry date and the posted entry id. from/to filter on the record date; savedViewId applies a G00 saved view.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true in annotations, the description appropriately adds behavioral detail beyond that: explains row structure (export vs posting), fields per type, and newest-first ordering. It does not cover pagination or limits, but the read-only nature is covered by annotations and the row description is helpful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence that front-loads the main purpose and then efficiently details row types and filters. No fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list operation with no output schema, the description adequately explains the return content (row types and fields), ordering, and filter semantics. It does not mention pagination, but this is likely not critical for a history list and the description is otherwise complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description is the only source. It explains 'from/to' as date filters and 'savedViewId' as applying a G00 saved view. The required workspaceId is not explicitly described, but it is a standard workspace identifier and the description covers three of four parameters meaningfully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: returns the payroll hand-off history, a union of exports and wage-journal postings, with row types and ordering. This is a specific verb-resource pairing that distinguishes it from sibling tools like export_journal or wage_journal_post, even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides usage context for the filters (from/to on record date, savedViewId for G00 saved views) but does not explicitly state when to use this tool over alternatives. The union nature implies it's the comprehensive view, but no exclusion guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_period_locksARead-only
List all period locks for the workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is known. The description adds the scope qualifier 'all ... for the workspace', but reveals nothing else about response behavior, ordering, or potential limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with no filler. Every word contributes to understanding the tool's function and scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read-only list, the description is mostly adequate, especially with the readOnlyHint annotation. However, without an output schema, it does not describe what a period lock is or what data the response will contain, which could leave some ambiguity for an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage for the single parameter, so the description must compensate. The phrase 'for the workspace' connects workspaceId to the returned locks, but no additional meaning, format, or semantics are provided beyond the parameter name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'List' with a clear resource, 'all period locks', and states workspace scoping. This clearly distinguishes it from sibling mutation tools like lock_period and unlock_period.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context implies that this tool is for viewing period locks, but it does not explicitly state when to use it versus alternatives, nor does it mention when not to use it. No exclusions or alternative routing are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_pluginsARead-only
Die installierten Erweiterungen (P5, US-G02.2/3/4): every plugin manifest for the workspace newest first, each with its status, capability count, requested/granted scopes and compat verdict. savedViewId applies a stored G00 view over the plugin kind (filter by status or capability kind); its filters merge underneath any status named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already communicates that this is a safe read operation. The description adds valuable behavioral details beyond that: the result includes every plugin manifest, is sorted newest first, and includes the plugin status, capability count, scopes, and compatibility verdict. It also discloses how savedViewId interacts with an explicitly passed status, which gives the agent a clear picture of behavior without contradicting the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately sized but contains a German preamble referencing 'P5, US-G02.2/3/4' that adds no operational value for an AI agent. The key information (returns all plugin manifests with specific fields, sorted newest first) is present, and the savedViewId behavior is explained, but the structure could be tighter by removing the requirement-ticket reference and better separating the output semantics from filter behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, the description reasonably enumerates the returned fields and ordering, which is helpful. However, it leaves the required workspaceId and the status parameter values unspecified, and it does not mention pagination limits, if any, or what a 'G00 view' means. For a list tool with three parameters and no output schema, this is incomplete but not severely lacking.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description bears the burden of explaining parameters. It only meaningfully explains savedViewId (applies a stored view, filters merge under an explicit status). It does not explain the required workspaceId or the status parameter beyond a vague reference to 'filter by status or capability kind' embedded in the savedViewId explanation. Two of three parameters remain undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly specifies the verb (list) and resource (installed plugins/Erweiterungen) for the workspace, and enumerates the returned fields: status, capability count, requested/granted scopes, and compat verdict. The phrase 'installierte Erweiterungen' implicitly distinguishes this from registry-oriented siblings like search_plugin_registry and get_plugin_registry_entry, which operate on a plugin catalog rather than installed workspace plugins.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when the tool is appropriate by focusing on installed plugins, but it does not explicitly state when to prefer this over siblings such as search_plugin_registry or get_plugin. It does provide some usage guidance for savedViewId, explaining how its filters merge with an explicit status, but this is about parameter usage, not tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_reconciliationARead-only
The reconciliation board: matched, unmatched and partial txns. A single-statement call (statementId) also returns reconciled (true once the ledger bank account matches the statement's closing balance to the Rappen, D64 as-of; absent when the Bankkonto is foreign-currency or the message carried no balance). A cross-statement call (bankAccountId/from/to, or savedViewId) omits it and returns the filtered lists only.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| savedViewId | No | ||
| statementId | No | ||
| workspaceId | Yes | ||
| bankAccountId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true, so the safe-read nature is already known; the description adds meaningful behavioral detail by explaining the conditional presence of the 'reconciled' field and its exact semantics (ledger bank account matching the statement closing balance to the Rappen, D64 as-of, omitted under specified conditions). It also clarifies that cross-statement calls omit reconciled, adding nuance beyond the annotation without contradicting the read-only hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient, front-loading the core purpose first and then covering the two call modes in one compound passage. No words are wasted, though some phrasing ('Bankkonto', 'D64 as-of') is jargon-heavy and could be more accessible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only listing tool with no output schema, the description covers the main call patterns, required context, and output differences between them. It doesn't describe pagination, list shape, or the meaning of 'partial txns' in detail, but the critical usage distinctions are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: statementId, bankAccountId, from/to, and savedViewId all receive functional meaning tied to call modes. workspaceId is not explicitly explained, but it's required and reasonably inferable. The parameter semantics go well beyond the bare schema, though format details for dates are not given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific resource—'the reconciliation board'—and states what it returns: matched, unmatched, and partial transactions. It further distinguishes single-statement vs cross-statement call behavior, which helps define the tool's scope. It doesn't explicitly differentiate from sibling tools like list_bank_statements or review_bank_txn, but the resource and variants are clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete guidance on which parameters select which mode: statementId triggers a single-statement call, while bankAccountId/from/to or savedViewId triggers a cross-statement call. This is actionable context for when to use each variant, though it does not explicitly state when to prefer this tool over related reconciliation or bank-statement tools, nor does it mention exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_recurring_schedulesARead-only
List the Serienrechnungen with their cadence, next run, occurrence count, status and the LAST run outcome (so a permanently failing schedule is visible on the list, not only in its history), newest first, filtered by status or contact. savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| contactId | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description adds behavioral context: it discloses the 'newest first' ordering, that the last run outcome is included to surface failing schedules, and how savedViewId filters merge with explicit ones. This is useful beyond annotations, though it omits pagination behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose and key fields, and includes the saved view behavior without redundancy. Every sentence adds value, and the structure is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with no output schema, the description covers the essential details: what is listed, ordering, filters, and saved view behavior. It lacks explicit pagination information, which is a minor gap, but overall it provides sufficient guidance for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by explaining status and contactId as filters and detailing savedViewId merging. workspaceId is not explained but is self-evident as a required workspace scope. The description adds meaning to the parameters, though it does not enumerate possible status values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists recurring schedules (Serienrechnungen) and enumerates the exact fields returned: cadence, next run, occurrence count, status, and last run outcome. It also specifies ordering and filters, making the purpose unambiguous and distinguishable from sibling get_recurring_schedule.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool (to list schedules with filters) and explains how savedViewId merges filters. It does not explicitly contrast with get_recurring_schedule for single-schedule retrieval, but the name and list semantics are clear. No explicit exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_restorable_backupsARead-only
List the .tillbackup bundles in this machines backup directory (the support dirs backups/, or TILL_BACKUP_DIR), newest first, with each one`s creation time, schema generation, entry count and whether this runtime can restore it, plus the generation this runtime expects. Pre-workspace: the restore door on a fresh install reads it. Lists no exports (an export is not a restore source) and verifies nothing (verify_backup does).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true, so the read-only nature is known. The description adds meaningful behavioral context beyond that: it does not verify backups, does not list exports, and includes runtime-specific compatibility information ('whether this runtime can restore it' and 'the generation this runtime expects'). It also explains the source directory including the TILL_BACKUP_DIR override.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compactly written in three sentences, with the primary purpose front-loaded, followed by return details, then usage context and exclusions. Every sentence provides distinct value with no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description fully explains what the tool returns: creation time, schema generation, entry count, restore compatibility, and expected generation. It also covers when it is used and what it deliberately omits. This is complete for a zero-parameter, read-only list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description confirms no inputs are needed by referring only to machine environment (backup directory), which is sufficient. There are no parameter schemas to elaborate on.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and resource ('.tillbackup bundles in this machine's backup directory') with concrete details: newest first, each bundle's creation time, schema generation, entry count, and restore compatibility. It clearly differentiates itself from sibling tools like list_backups and verify_backup by specifying its exact scope and what it returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states when this tool is used ('Pre-workspace: the restore door on a fresh install reads it') and explicitly contrasts with alternatives: 'Lists no exports (an export is not a restore source) and verifies nothing (verify_backup does)'. This gives an agent clear conditions for when to call this tool versus others.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_rolesARead-only
List every assignable role with its resolved capability bundle and its member count, plus the capability registry itself so a client can render the whole matrix.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description aligns with that by describing a read-only listing operation. It adds useful details about the returned data (capability bundles, member counts, registry) without contradicting the safety profile, though it omits potential pagination or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that efficiently communicates the core action and return details. It is concise without unnecessary filler, though it could benefit from splitting the long sentence for clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with readOnly annotations, the description covers the primary output but lacks mention of pagination, size limits, or any preconditions (e.g., workspace existence). It is adequate for basic understanding but not fully complete for an agent handling large role sets.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has a single required parameter, workspaceId, with no description (0% coverage). The tool description does not mention this parameter or explain its role, leaving the agent to infer that it scopes the listing to a workspace. This is a notable gap given the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists every assignable role, along with its capability bundle, member count, and the capability registry. This specific verb-resource pairing distinguishes it from role mutation tools like define_role, archive_role, and set_role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for rendering the role-capability matrix, but it does not explicitly state when to use this tool versus alternatives. No exclusions or comparisons to sibling list tools (e.g., list_members) are provided, leaving the selection reasoning to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_saved_viewsARead-only
List the saved views available for an entity kind: this session's own first, then the workspace-shared ones. Another session's personal views are never returned.
| Name | Required | Description | Default |
|---|---|---|---|
| entityKind | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds meaningful behavioral context: the ordering (session's own first, then workspace-shared) and an explicit exclusion (another session's personal views). This goes beyond what annotations convey, though it does not cover details like pagination or response format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The core action is front-loaded, and the scoping rules are stated in a compact, readable way.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with two parameters and no output schema, the description provides the key behavioral semantics (what is returned, ordering, and exclusions). The main missing piece is the meaning of workspaceId, which prevents full standalone usability.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for entityKind and workspaceId. It clarifies entityKind's role ('for an entity kind'), but it never mentions workspaceId or explains its purpose or constraints. This is a meaningful gap for an agent needing to construct correct calls.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'List the saved views available for an entity kind.' It also adds distinguishing details about the scope and ordering (session's own first, then workspace-shared), which separates it from other saved-view tools like create_saved_view or update_saved_view.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly communicates when to use this tool (when you need to list saved views for an entity kind) and excludes a specific case: another session's personal views are never returned. However, it does not explicitly name alternatives or say when not to use it, so it falls short of the strongest routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_unmatched_incomingARead-only
The Abgleich queue: every registered incoming credit with its LIVE score (an open row re-scores on read, so a later-issued invoice is found), the proposed invoice with the exact Rappen figures (open amount, unpaid Mahngebühr, delta), the applied rows with their payment link and correction chain, the counts per lane, and the auto-apply dial state. Filter by Bankkonto, status (open/applied/dismissed) or value-date range. savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here. Reads only.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| status | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| bankAccountId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, it discloses important runtime behavior: open rows re-score on read, exact Rappen-level figures are returned, and savedViewId merges stored filters underneath explicit filters. This is meaningful operational context the annotation alone does not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded, starting with the core queue identity and covering output, filters, and saved-view behavior. It is one long run-on paragraph and could be better structured, but every clause adds relevant detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list operation with no output schema, the description covers the returned components, supported filters, and saved-view behavior in detail. The agent has enough context to call the tool correctly, with only minor gaps like exact date formats.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description compensates well by explaining bankAccountId as Bankkonto, listing status values open/applied/dismissed, describing from/to as a value-date range, and detailing savedViewId merge semantics. The only parameter left undocumented in both the schema and description is workspaceId.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific verb and resource: listing the Abgleich queue of unmatched incoming credits. It enumerates the returned fields and filters, which strongly distinguishes it from sibling reconciliation and matching tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the tool's context unmistakable: reviewing unmatched incoming credits with filters by bank account, status, or date range, and it explicitly says 'Reads only.' It does not explicitly name alternatives or when-not-to-use cases, but the context is clear enough for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_vendor_billsARead-only
Die Kreditoren: every vendor bill with its open amount, due date, days overdue and aging bucket, plus the tie-back to account 2000. reconciled compares two independent derivations (the bills and their allocations against the posted balance of 2000), so a movement A17 does not model shows up as a stated difference rather than as a silently wrong total. status filters the lifecycle (draft/posted/void); settlementStatus filters the derived payment state (unpaid/partly_paid/paid). savedViewId applies a G00 saved view underneath the explicit filters.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| status | No | ||
| vendorId | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| settlementStatus | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the readOnlyHint annotation by explaining how the reconciled field compares two independent derivations and how unmatched movements surface as stated differences. It also clarifies the semantics of status, settlementStatus, and savedViewId, giving the agent a clear picture of the tool's filtering behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient: the first sentence front-loads the core purpose, the second explains the reconciliation behavior, and the third details the filters. No words are wasted, though 'Die Kreditoren' adds mild ambiguity for English-only readers.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the key behavioral aspects (reconciliation, filter semantics) and the list output, but does not document the date range parameters (to/from) or vendorId, which are still in the schema without any description. Given the absence of an output schema and the tool's complexity, these omissions leave some gaps in what an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description compensates by explaining status, settlementStatus, and savedViewId, but leaves to, from, and vendorId semantically implicit. An agent would need to infer that to/from are likely date range bounds and vendorId is an ID filter, which is a notable gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as listing vendor bills with open amount, due date, days overdue, aging bucket, and tie-back to account 2000. It specifies the resource and the fields returned, but does not explicitly name or differentiate it from closely related siblings like list_payable or get_vendor_bill.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the description's focus on vendor bills with aging and reconciliation to account 2000. There is no explicit guidance on when to prefer this tool over alternatives (e.g., list_payable for payables or get_vendor_bill for a single bill), nor any exclusionary conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_workspacesARead-only
List the workspaces in this database this session may open (id, name, currency, fiscal year start), so a client can re-open one. Archived mandates are hidden unless includeArchived is true; a provisioned workspace is listed only for its accepted members. savedViewId applies a saved roster view (G00): pass savedViewWorkspaceId (the book the view is stored in) beside it; stored filters merge underneath any named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| savedViewId | No | ||
| includeArchived | No | ||
| savedViewWorkspaceId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses default behavior (archived mandates hidden unless includeArchived is true), membership filtering for provisioned workspaces, and the saved-view merge semantics. With readOnlyHint=true, this adds substantial context beyond the annotation without contradicting it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then compact behavioral details and parameter semantics. Each clause earns its place, though it is dense and uses domain-specific shorthand that could be clearer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the key behaviors, return fields, and parameter semantics. With no output schema it could be even more complete, but pagination or sorting details are not critical for a listing tool; overall it is sufficient for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but the description explains all three parameters: includeArchived toggles archived visibility, savedViewId applies a saved roster view, savedViewWorkspaceId provides the book where the view is stored. Some jargon remains (G00, 'roster view'), so it isn't fully self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clear verb+resource: 'List the workspaces in this database this session may open' plus the return fields, and the client-reopen scenario distinguishes it from sibling workspace tools. It is immediately distinct from get_workspace vs env_list vs archive_workspace.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear scenario ('so a client can re-open one') and explains when to pass includeArchived and the savedView parameters. Does not explicitly name alternative tools or state when not to use them, so it lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
location_archiveA
Soft-archive a location (and its empty descendants). Refused with location_has_stock when the location or any descendant still has non-zero on-hand, location_in_use when an open stocktake references it, or cannot_archive_default when it is a default. Deletion is never offered; historical movements keep their location_id.
| Name | Required | Description | Default |
|---|---|---|---|
| locationId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses that the operation is a soft-archive (non-destructive), affects empty descendants, refuses under three specific conditions with named error codes, and preserves historical movements. This is rich context beyond the schema. It does not mention idempotency behavior or response format, but the key side effects are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler. The purpose is front-loaded, followed by refusal conditions and a note on non-deletion. Every sentence adds value and the structure is clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description explains the action's purpose, conditions, and side effects. It does not describe the success response, but for an action tool that is often minimal. The notion of 'empty descendants' is clarified by the refusal condition about non-zero on-hand. Overall, an agent can call it correctly with this description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not explain any of the three parameters. Since schema coverage is 0%, the description must compensate, but it only mentions location and descendants without clarifying the idempotencyKey or workspaceId. While the parameter names are self-explanatory, no additional meaning is added, leaving the agent to infer from names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('soft-archive') and resource ('a location') with an explicit scope ('and its empty descendants'). It distinguishes this from other location tools (create, update, list, get, tree) and from delete tools by clarifying it's non-destructive. The refusal conditions also sharpen the purpose by describing what it cannot do.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (soft-archive as an alternative to deletion) by stating 'Deletion is never offered; historical movements keep their location_id.' However, it does not explicitly name alternatives or contrast with other archive/delete tools. The refusal conditions are about edge cases, not about choosing this tool over another.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
location_createA
Create a location under a warehouse. code is 1-30 chars, unique within the warehouse case-insensitively (duplicate_code). parentId nests it (null = root under the warehouse); the parent must belong to the SAME warehouse (parent_warehouse_mismatch) and the depth is capped at 8 (max_depth_exceeded). locationType is one of zone|aisle|shelf|bin|staging|other (invalid_location_type). The FIRST location in a warehouse becomes its default; pass isDefaultForWarehouse to force it.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | ||
| name | Yes | ||
| parentId | No | ||
| description | No | ||
| warehouseId | Yes | ||
| workspaceId | Yes | ||
| locationType | No | ||
| idempotencyKey | Yes | ||
| isDefaultForWarehouse | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well. It discloses validation rules (code length/uniqueness, parent warehouse match, depth cap), enumerates locationType allowed values, and reveals the side effect that the first location becomes the default unless overridden. It even lists potential error codes (duplicate_code, parent_warehouse_mismatch, max_depth_exceeded, invalid_location_type), giving the agent a complete behavioral picture.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence that packs a lot of information without fluff. It front-loads the action and then enumerates constraints logically. Though long, it is concise in the sense that every clause adds value; it could be broken into bullets for readability, but the structure is acceptable and not verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lacking an output schema and having zero parameter descriptions, the description covers the essential context: what the tool does, key validation rules, allowed values, side effects, and error scenarios. It gives the agent enough to invoke correctly without needing to infer hidden business logic, making it a well-rounded and complete description for this complexity level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains code, parentId, locationType, and isDefaultForWarehouse with meaningful detail. However, it does not elaborate on workspaceId, warehouseId, name, description, or idempotencyKey. While some (warehouseId) are obvious from the purpose, others like idempotencyKey and name rely on the agent's general knowledge. It covers the most critical parameters but not all, earning a 4 in the absence of schema help.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Create' and the resource 'location under a warehouse', and distinguishes it from siblings like location_update or location_list by focusing on creation semantics. It also names key constraints that define what this tool does uniquely.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose is unambiguous: creating a location in a warehouse. It does not explicitly contrast with alternatives, but the constraints (e.g., parent warehouse matching, depth cap) and the existence of sibling tools like location_update, location_archive make the creation intent obvious. It lacks explicit when-not-to-use guidance, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
location_getARead-only
Read one location by id, archived or not, including its warehouse, parent, materialised path and depth.
| Name | Required | Description | Default |
|---|---|---|---|
| locationId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already signals the operation is non-mutating; the description adds value by specifying it returns archived-or-not locations and the included hierarchy fields (warehouse, parent, path, depth). No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One concise sentence that packs the verb, scope, and return fields with no filler. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has only two parameters, none nested, and no output schema. The description covers the read behavior, the included fields, and the archived-or-not scope. It does not state error conditions or what happens when the location is missing, but for a simple read with readOnlyHint, the description is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the resource (location) and that locationId identifies the record, but it does not explain the role of workspaceId beyond the schema's required field. Some meaning is added, but not enough to fully bridge the lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read'), a singular resource ('one location by id'), and clarifies scope ('archived or not'). It enumerates the returned data (warehouse, parent, materialised path, depth), which distinguishes it from sibling list/tree/upsert location tools without needing their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the right use case: fetching a single, specific location including archived ones and related hierarchy data. It does not explicitly contrast with siblings like location_list or location_tree, so it loses a point for no explicit when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
location_listARead-only
List locations (flat), filterable by warehouseId, parentId (null = roots), active, and a case-insensitive search over code and name. includeDescendants=true with a parentId returns the whole subtree. Accepts a savedViewId (G00 saved-view seam).
| Name | Required | Description | Default |
|---|---|---|---|
| active | No | ||
| search | No | ||
| parentId | No | ||
| savedViewId | No | ||
| warehouseId | No | ||
| workspaceId | Yes | ||
| includeDescendants | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true, which already establishes the safety profile. The description adds genuinely useful behavior beyond that: parentId null = roots, case-insensitive search scoped to code and name, and includeDescendants returning the whole subtree. The cryptic "(G00 saved-view seam)" reference is opaque jargon but still discloses the saved-view acceptance behavior. No contradiction with the read-only annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core verb+resource, and every sentence adds parametric or behavioral value. The only slight blemish is the cryptic "G00 saved-view seam" parenthetical, which trades clarity for terseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter list tool with no output schema, the description covers the query-side semantics nearly completely — filters, subtree behavior, and saved-view support. Gaps are minor: the return shape and pagination are not mentioned, and the G00 reference is unexplained. With readOnlyHint already covering the safety profile, this is close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the full burden and compensates well: it adds semantics for 6 of 7 parameters — warehouseId, parentId (null = roots), active, search (case-insensitive over code/name), includeDescendants, and savedViewId. Only the required workspaceId is left implicit, which is acceptable as a standard workspace-scope parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"List locations (flat)" gives a specific verb and resource, and the parenthetical "flat" directly distinguishes it from the sibling location_tree tool. The named filters (warehouseId, parentId, active, search) further pin down what this listing endpoint is for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly conveys the filtering context and the includeDescendants=true + parentId subtree path, which implies when this tool can substitute for a tree view. However, it never explicitly names alternatives (e.g., location_get for a single record, location_tree for a native hierarchy) or states when NOT to use it, leaving the boundary between location_list and location_tree to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
location_set_defaultA
Make this location the default for its warehouse, demoting the previous default in the same transaction. An archived location is refused with location_archived.
| Name | Required | Description | Default |
|---|---|---|---|
| locationId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses a transactional behavior (demoting the previous default in the same transaction) and an error condition (location_archived for archived locations). This is meaningful behavioral context beyond what the schema shows. It doesn't mention permissions or reversibility, but the core side effects are clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The core action and the key side effect are front-loaded, and the error condition is stated in the second sentence. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple mutation tool with 3 required parameters and no output schema, the description covers the action, the transactional side effect, and a key error case. It doesn't explain the return value or idempotency semantics, but those are less critical for a straightforward set-default operation. The sibling list is large, but the description's specificity is enough to route an agent correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the role of locationId implicitly ('this location') and workspaceId implicitly ('its warehouse'), but it doesn't explicitly map each parameter. idempotencyKey is not mentioned at all. The description adds some meaning but leaves the agent to infer parameter roles from the schema names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Make this location the default'), a specific resource ('for its warehouse'), and a key behavioral detail (demoting the previous default in the same transaction). This clearly distinguishes it from sibling tools like location_create, location_update, and warehouse_set_default.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when you want to set a location as the default for its warehouse. It also mentions an exclusion (archived locations are refused), which helps an agent avoid calling it in the wrong situation. However, it doesn't explicitly name alternative tools or state when NOT to use it beyond the archived-location case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
location_treeARead-only
Return the nested location tree for one warehouse (roots with recursive children), for the GUI tree view.
| Name | Required | Description | Default |
|---|---|---|---|
| warehouseId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true already covering the safety profile, the description adds useful behavioral detail: the result is a nested tree with roots and recursive children, scoped to one warehouse. It does not contradict the annotation, and the added shape information is valuable beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no redundant words. It front-loads the primary action and scope, and the GUI context is appended naturally. Every phrase contributes to understanding the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tree-fetching tool, the description conveys the core behavior and intended display context. It is not fully complete because workspaceId semantics are left to inference and there is no output schema, but the tool is simple enough that an agent can likely call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only mentions 'one warehouse', which loosely maps to warehouseId, but it does not explain workspaceId or clarify the relationship between the two required parameters. For a two-parameter tool with no schema descriptions, this is a meaningful gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Return'), a specific resource ('nested location tree'), and a scope ('one warehouse'). 'For the GUI tree view' adds a concrete purpose, and the mention of 'nested' and 'recursive children' distinguishes it from a flat location_list or single location_get tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use it: when a tree-structured view of locations for a single warehouse is needed. It does not explicitly name alternatives or conditions when not to use it, but the GUI tree view context is specific enough to guide selection among the many location-related siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
location_updateA
Edit a location through a patch object (name, description, locationType, parentId). Re-parenting via parentId is allowed only within the same warehouse (parent_warehouse_mismatch) and never into the location itself or a descendant (location_cycle); the whole subtree moves and its materialised paths and depths are rewritten. Only the fields present in patch change.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | No | ||
| locationId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It excellently describes re-parenting constraints (same warehouse, no cycles), the subtree move and rewriting of materialised paths/depths, and the partial-update behavior (only fields present in patch change). This is rich, actionable information beyond what any schema could convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose, and every sentence adds distinct value. It avoids redundancy and conveys the most critical information efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the nested patch object and no output schema, the description covers the most complex aspects (re-parenting semantics, subtree rewriting). However, it does not explain required parameters like workspaceId, locationId, or idempotencyKey, nor describe the return value. These are likely standard across sibling tools, but for a standalone evaluation, some information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explicitly lists the patch fields and gives deep semantics for parentId, including error codes (parent_warehouse_mismatch, location_cycle) and subtree behavior. Other fields (name, description, locationType) are only named, but their meaning is self-evident. This is a solid compensation for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Edit' and the resource 'location', and enumerates the patch fields (name, description, locationType, parentId). This unambiguously identifies the tool's purpose and differentiates it from sibling location tools (create, archive, list, get, tree) through the edit action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for editing an existing location but does not explicitly state when to use this tool versus alternatives (e.g., location_create for new locations). There is no mention of exclusions or conditions that would select a different sibling tool. The guidance is implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lock_periodA
Lock a month or year, soft or hard. CONSEQUENCE: Locks the period against further posting; a hard lock seals it permanently and cannot be unlocked.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | ||
| period | Yes | ||
| reason | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral transparency burden. It explicitly discloses the key consequence: locking prevents further posting, and a hard lock is permanent and irreversible. This is strong, useful behavioral context beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loads the core action, and uses a clear consequence statement. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the central behavior and consequence well, but the lack of any parameter explanations, return information, or guidance about soft locks being reversible leaves gaps for an agent with no annotations and no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only partially compensates. It maps 'month or year' to period and 'soft or hard' to kind, but it does not explain workspaceId, idempotencyKey, or reason, leaving required parameters underspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: lock a month or year, either soft or hard, against further posting. It distinguishes itself from the sibling unlock_period by explicitly saying a hard lock cannot be unlocked.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose implies when to use it: when a period should no longer accept postings. However, it does not explicitly mention alternatives like close_month, close_year, or unlock_period, nor does it state when a soft lock should be preferred over a hard lock.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lot_archiveA
Soft-archive a lot (status archived). Refused with lot_has_balance while its derived on-hand is non-zero. Deletion is never offered; the movements that reference it keep their lot_id.
| Name | Required | Description | Default |
|---|---|---|---|
| lotId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the archive is soft, sets status to archived, refuses on non-zero derived on-hand, and preserves movement references. It stops short of reversibility or authorization, but is transparent about the core side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two crisp sentences front-load the action and then add the key constraints and referential behavior. No wasted words; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple archive mutation with no output schema, the description covers the main behavior, refusal condition, and referential integrity. It doesn't discuss return/error format or required permissions, but that's acceptable for a tool of this simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
All three parameters have no schema-level description (0% coverage), and the description never mentions workspaceId, lotId, or idempotencyKey. For example, idempotencyKey's role is unexplained. The description does not compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb+resource ('Soft-archive a lot') and clarifies the exact effect ('status archived'). It also notes that deletion is never offered, which distinguishes it from a harder removal. Sibling tools aren't named, so it doesn't reach a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to call it (to archive a lot) but does not contrast it with `lot_set_status` or other lot operations, nor does it explicitly state when to use it versus alternatives. The refusal condition is a constraint, not usage guidance, so it falls at 'implied usage'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lot_createA
Create a lot (a batch) against a lot-tracked item. number is 1-60 chars, unique per item case-insensitively (lot_number_taken); the item must be lot-tracked (else tracking_not_applicable). Optional expiryDate, manufacturedDate, supplierReference, notes and an initial status (default open). No quantity column: on-hand is always derived from movements.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| itemId | Yes | ||
| number | Yes | ||
| status | No | ||
| expiryDate | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| manufacturedDate | No | ||
| supplierReference | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the behavioral disclosure burden. It reveals important non-obvious behaviors: number uniqueness rules, error conditions like lot_number_taken and tracking_not_applicable, the default status, and the fact that on-hand quantity is derived from movements. It does not mention idempotency behavior or permissions, but the key behavioral traps are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three dense sentences with no filler. It front-loads the core action, then efficiently covers preconditions, optional parameters, and the critical quantity caveat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter tool with no output schema and no annotations, this description covers the non-obvious inventory semantics and required preconditions well. The main omissions are idempotencyKey behavior and any indication of return value, but an agent still has enough context to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining parameters. It adds real meaning: number length and case-insensitive uniqueness, optional fields, and default status. workspaceId and idempotencyKey are not explained beyond the schema, so the compensation is strong but not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and object: 'Create a lot (a batch) against a lot-tracked item.' It clearly scopes the operation and is distinct from siblings like lot_update, lot_set_status, and lot_archive because it is explicitly the creation operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: when creating a lot for a lot-tracked item, and it explicitly states the precondition that the item must be lot-tracked. It does not name alternative tools such as serial_create or lot_set_status, so it stops short of full when-not/alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lot_getARead-only
Read one lot by id, plus its derived on-hand (the live SUM of movements for the lot).
| Name | Required | Description | Default |
|---|---|---|---|
| lotId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, so safety is covered. The description adds valuable behavioral context: that on-hand is not a stored value but a derived live SUM of movements. This goes beyond the annotation and helps the agent understand the data semantics and potential cost (aggregation). No contradiction found.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, front-loaded with the primary action and then the key derived detail. There is no fluff. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with two required params and a readOnlyHint, the description covers the essential purpose and the notable derived value. It doesn't describe return shape or error cases, but with no output schema and a clean read operation this is not a critical gap. It is complete enough to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. It explains that the tool reads 'by id', which maps to lotId, but says nothing about workspaceId. Neither parameter has a schema descriptionheb, and the description does not clarify the role or format of workspaceId. The agent is left to infer that workspaceId is a scope qualifier.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read') and a specific resource ('one lot by id'), and adds a distinctive detail ('derived on-hand' with live SUM). This clearly differentiates it from siblings like lot_list and lot_search. The agent understands exactly what the tool does and for what resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when a specific lotId is known and on-hand is needed, but it does not explicitly say when to prefer this over alternatives like lot_search or inventory_on_hand_by_lot. No alternatives are named, and no when-not-to-use guidance is given. Usage is implied but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lot_listARead-only
List lots, filterable by itemId, status, an expiryBefore cutoff (FEFO), and a case-insensitive search over number and supplierReference. Archived lots are hidden unless includeArchived or an explicit status is given. Each row carries its derived on-hand. Accepts a savedViewId (G00 saved-view seam).
| Name | Required | Description | Default |
|---|---|---|---|
| itemId | No | ||
| search | No | ||
| status | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| expiryBefore | No | ||
| includeArchived | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations only declare readOnlyHint=true, which the description does not contradict. The description adds meaningful behavioral context beyond the annotation: archived lots are hidden unless includeArchived or an explicit status is given, and each row carries its derived on-hand. These are useful behaviors not present in the schema or annotations. It doesn't cover pagination or ordering, but the readOnly annotation lowers the burden and the added behavior is substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the core action and filters, with the archived behavior and savedViewId noted briefly. No wasted words; every clause adds information. The structure is easy to scan without being terse to the point of omitting key behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's relative complexity (7 parameters, no output schema, no enums), the description covers the main call semantics, filtering behavior, archived handling, and the saved-view seam. It does not mention output shape or pagination, though the absence of an output schema means those are under-specified; the core selection and invocation semantics are complete enough for an agent to use it correctly for typical list-and-filter purposes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the bare input schema. The description adds semantics for itemId (filter), status (also affects archived visibility), expiryBefore (FEFO cutoff), search (case-insensitive over number and supplierReference), includeArchived (overrides default hiding), and savedViewId (accepted as a saved-view seam). It doesn't explain expiryBefore's date format or the valid status values, which is a minor gap, hence 4 rather than 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and resource ('lots'), enumerates the filter dimensions (itemId, status, expiryBefore/FEFO, case-insensitive search), and distinguishes behavior from siblings like lot_search by noting the derived on-hand and archived-lot handling. The scope is clear and an agent can differentiate it from the related lot_search sibling without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool: listing lots with filters, including the special FEFO cutoff and case-insensitive search. It notes the archived-lot default behavior and the includeArchived override. It does not explicitly name alternative tools (e.g., lot_search or inventory_on_hand_by_lot) for when a different tool should be chosen, so it earns a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lot_searchARead-only
Full-text lot search over number and supplierReference (case-insensitive), capped at 100 rows. Empty query returns an empty list.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds useful behavioral details beyond that: case-insensitive matching, a hard 100-row cap, and empty-query behavior. This helps the agent predict results without seeing output.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences deliver scope, behavior, and an important edge case with no filler. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with a simple two-parameter schema, the description covers the core behavior: what is searched, case sensitivity, result cap, and empty-query semantics. It does not describe return structure or sorting, but those are not critical for invocation given the simple schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies that the query targets number and supplierReference, which adds meaning to the 'query' parameter, but it does not explain the workspaceId parameter or query syntax more deeply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (full-text search), the target resource (lots), and the exact fields searched (number and supplierReference). This distinguishes it clearly from sibling tools like lot_list, lot_get, and serial_search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for when to use the tool: when performing free-text, case-insensitive lookup on lot identifying fields. It does not explicitly name alternatives or exclusions, but the use case is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lot_set_statusA
Change a lot status: open | held | expired | closed | archived, with an optional reason. Setting closed or archived while the lot still has non-zero derived on-hand is refused with lot_has_balance (quarantine or expire it instead, which keeps the quantity). Never mints a stock movement.
| Name | Required | Description | Default |
|---|---|---|---|
| lotId | Yes | ||
| reason | No | ||
| status | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses the refusal condition (lot_has_balance), explains that closing/archiving is blocked when balance exists, and explicitly states that the operation never mints a stock movement. It could also mention idempotency semantics given the required idempotencyKey, but the disclosed behaviors are substantive and useful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the action and allowed values, followed by the critical edge-case rule. Every sentence adds value and there is no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is largely complete for invocation: it defines the allowed statuses, notes the optional reason, explains a key refusal with an alternative, and clarifies that no stock movement is created. The main omission is the success/return shape, but that is not essential for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does add meaning for the 'status' parameter by listing allowed values and indicates that 'reason' is optional, but it does not explain workspaceId, lotId, or idempotencyKey beyond their names. This is partial compensation only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Change a lot status') and the resource ('lot'), and enumerates the exact allowed status values: open, held, expired, closed, archived. This distinguishes it from related sibling tools like lot_archive or inventory_move because it is specifically a status transition operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides concrete guidance for an important edge case: if the lot has non-zero derived on-hand, setting closed or archived is refused, and the description explicitly suggests using quarantine or expire instead. It does not broadly contrast with sibling tools like lot_archive or serial_set_status, but the condition-specific alternative is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
lot_updateA
Edit a lot through a descriptive patch (number if still unique, expiryDate, manufacturedDate, supplierReference, notes). The status is changed via lot_set_status, not here. Only the fields present in patch change.
| Name | Required | Description | Default |
|---|---|---|---|
| lotId | Yes | ||
| patch | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral disclosure burden. It discloses partial-update semantics ('Only the fields present in patch change'), a uniqueness constraint ('number if still unique'), and the boundary that status is not modified here. It does not cover idempotency key behavior or error cases, but the core mutation behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no fluff. The action verb and resource appear first, followed by a compact parenthetical field list and two crucial behavioral clarifications. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, no annotations, and a nested patch object, the description covers the main edit scope and status boundary but leaves gaps around idempotency key purpose, what the tool returns, and how uniqueness failures are handled. These gaps are noticeable, but the core usage is not misleading.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaning to the patch fields by listing them and noting the uniqueness caveat for the number field, which goes beyond the schema types. However, it does not explain the required top-level parameters workspaceId, lotId, or idempotencyKey, and the schema description coverage is 0%, so the description only partially compensates for the missing parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool edits a lot via a patch, enumerates the editable fields, and explicitly distinguishes itself from lot_set_status for status changes. This gives a specific verb+resource and distinguishes it from the most relevant sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says status changes are handled by lot_set_status, not here, which is a clear when-not-to-use signal. It also clarifies that only fields present in the patch change, helping the agent understand partial updates. It does not explicitly name other alternatives like lot_create or lot_archive, but the status exclusion is strong guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_accounts_listARead-only
Die verbundenen Mail-Konten (P5): adapter, address, store path, enabled, and when each was last indexed, plus the supported adapter list for the connect affordance.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers the safety profile, and the description adds value by specifying the exact returned fields and the additional adapter list for the connect affordance. Nothing suggests side effects, and there is no contradiction with the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence that front-loads the resource and field list with minimal waste. The unexplained 'P5' slightly reduces clarity but does not make the description verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with one parameter and no output schema, the description covers the return fields and the extra adapter list. It does not explicitly state workspace scoping or pagination, but these are minor gaps for this simple tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention the workspaceId parameter or clarify that results are scoped to the given workspace. With only one parameter, the name is self-explanatory, but the description fails to compensate for the low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource (connected mail accounts) and enumerates the returned fields (adapter, address, store path, enabled, last indexed) plus the adapter list. However, it lacks an explicit verb like 'list' or 'returns,' and the unexplained 'P5' prevents a perfect score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for the connect affordance' gives an implicit usage context, and the tool name suggests listing mail accounts. But there is no explicit when-to-use guidance or comparison with alternatives like mail_connect or mail_threads_list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_connectA
Verbinde einen lokalen Mail-Speicher: point TILL at the store a mail client on this machine already writes (adapter apple_mail reads .emlx, thunderbird reads Maildir and mbox). Takes NO credential and never signs in anywhere: an unreadable path answers needs_mailstore, a readable path that matches no adapter answers unknown_mail_adapter. The same store under the same address stays ONE account.
| Name | Required | Description | Default |
|---|---|---|---|
| adapter | Yes | ||
| address | Yes | ||
| storePath | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses meaningful behavior: no credentials are used, unreadable paths produce needs_mailstore, unmatchable readable paths produce unknown_mail_adapter, and the same store/address stays one account. It does not describe success return shape, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: it states the operation, the local-source constraint, the no-login guarantee, and the two error paths in two sentences. Every sentence contributes behavioral or selection information without repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a connect tool with no output schema and no annotations, the description covers the main edge cases (unreadable path, unknown adapter) and the idempotency guarantee. The only gap is the lack of explicit success response shape, but the tool is simple enough that the description is largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds real meaning to adapter (apple_mail reads .emlx; thunderbird reads Maildir/mbox), storePath (local store location), and address (part of the account identity/idempotency rule). It does not explain workspaceId or idempotencyKey, but the most behaviorally significant parameters are well covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific operation ('Verbinde einen lokalen Mail-Speicher') and resource: the local mail store that a client already writes. It also distinguishes itself from the mail_* siblings by emphasizing that this is exclusively local, adapter-backed, and credential-free.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when to use the tool: when connecting to a local mail store (apple_mail, thunderbird) and explicitly says it never signs in anywhere, implying it is not for remote/online mail access. It stops short of naming an alternative tool, but the context is clear enough for an agent to route correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_drafts_listARead-only
Die von TILL zurückgeschriebenen Entwürfe (P5): locator, hash, draft-run provenance and reply target per draft, for one thread or across the workspace, newest first.
| Name | Required | Description | Default |
|---|---|---|---|
| threadId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true already declaring the safety profile, the description adds meaningful behavioral context: the output fields, scope flexibility, and sort order. It does not contradict annotations and helps an agent understand what to expect from the response. It could have mentioned pagination or empty-result behavior, but for a read-only list tool this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence that front-loads the key selector and packs all relevant details. There is no wasted text, but the unexplained acronyms (TILL, P5) make it less immediately readable. Overall it is concise but could be clearer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with a simple schema and no output schema, the description covers the core return fields and scope. However, it omits details like whether the list is paginated, how workspaceId is validated, and what 'draft-run provenance' means in practice. The jargon and sparse parameter documentation leave an agent with some uncertainty.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It maps the 'one thread or across the workspace' wording to the threadId and workspaceId parameters, giving them functional meaning. However, it does not explain parameter formats, requiredness, or the meaning of additionalProperties, so the compensation is only partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: listing drafts written back by TILL, with fields (locator, hash, provenance, reply target) and scope (thread or workspace). It distinguishes itself from mail_threads_list by focusing on TILL-generated drafts, though it never names the sibling. The internal references 'TILL' and 'P5' are unexplained, which slightly obscures the purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: to retrieve TILL-written drafts for a single thread or across the workspace. It implies the optional threadId vs workspaceId split, and the 'newest first' ordering informs expectations. It does not explicitly mention alternatives or exclusions, so it misses the top score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_draft_writeA
Schreibe einen Entwurf in den Entwurfsordner des Mail-Programms: renders an RFC-5322 reply (From the account address, To the counterpart, In-Reply-To the newest inbound message) into the local Drafts folder and persists locator + hash, not the body. The human reviews and sends it in their own mail app: TILL has no send verb and no SMTP anywhere (P8 by construction). One idempotency key writes exactly ONE Drafts message however often it is called; an unwritable folder answers drafts_not_writable.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | ||
| threadId | Yes | ||
| inReplyTo | No | ||
| draftRunId | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and delivers comprehensively. It discloses the side effect (writes to Drafts folder), the specific content rendered (RFC-5322 reply with From, To, In-Reply-To), that it persists locator+hash rather than body, the idempotency guarantee, and the error response drafts_not_writable. This is exemplary transparency for a write operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-rich, starting with the primary action in German then elaborating in English. Each sentence adds value: the purpose, the specific reply structure, the persistence behavior, the no-send constraint, idempotency, and error. It is slightly long but not bloated, and the front-loaded action ensures immediate clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (6 params, no output schema, no annotations), the description covers the purpose, behavior, idempotency, and error case, and touches on some parameter semantics. However, it leaves workspaceId and draftRunId unexplained, and does not describe the return value (e.g., whether it returns a locator/hash or just a success indicator). It also omits guidance on how to obtain required values like threadId. This is adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for six parameters. It explicitly explains idempotencyKey (idempotency), threadId (used to determine counterpart and inbound message), and body (the content). It also mentions that In-Reply-To is set to the newest inbound message, which relates to the inReplyTo parameter. However, workspaceId and draftRunId are not explained at all, leaving their purpose ambiguous. The description provides partial but not complete parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: write a draft into the mail program's Drafts folder. It specifies it renders an RFC-5322 reply with From/To/In-Reply-To and persists locator+hash, not the body. This distinguishes it from sibling tools like mail_drafts_list (reading) and mail_thread_get (viewing), and it is unambiguous about its write nature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly notes that TILL has no send verb and no SMTP, and that the human reviews and sends in their own mail app. This sets clear boundaries on when to use this tool (draft creation only, not sending). It also explains idempotency behavior (one key = one message) and the error case for unwritable folders. However, it does not explicitly compare to alternative tools like save_draft or mail_drafts_list, but its role is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_reindexA
Lies den Mail-Speicher neu ein: walk the store and derive the index (threads and messages keyed on Message-ID), storing a locator and a sha256 per message and NEVER a body (OP6 index-never-copy). Senders matching a contact e-mail resolve to that contact; unknown senders stay indexed with none. Messages deleted in the mail client leave the index (the store is the authority); unparseable ones are skipped and counted, never fatal. Answers { indexed, skipped, removed, skippedReasons }.
| Name | Required | Description | Default |
|---|---|---|---|
| accountId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so thoroughly: it discloses that bodies are never stored, deleted messages are removed, unparseable messages are skipped and counted without fatal errors, and sender-contact resolution behavior. These are precisely the non-obvious traits an agent needs to assess side effects and safety.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but appropriately sized, front-loading the core action and then covering invariants, edge cases, and the return shape. The bilingual opening and internal policy tag (OP6) are minor redundancies, but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides rich context about behavior and return shape, which is essential given no output schema and no annotations. It lacks parameter-level explanation and explicit usage timing, so it is not fully complete, but it is sufficient for an agent to understand what the tool does and what it returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not compensate by explaining workspaceId, accountId, or idempotencyKey. The tool's behavior implies account/workspace scoping, but the description never maps the parameters to the 'store' it references, nor explains idempotencyKey's role in retries.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('walk the store and derive the index') on a specific resource (the mail store), and names the index structure and key (threads/messages keyed on Message-ID). This makes the tool's purpose unambiguous and distinct from sibling mail tools like mail_threads_list or mail_draft_write.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The behavior described implies this is the maintenance/reindex operation for the mail store, but there is no explicit statement of when to use it versus alternatives, prerequisites, or exclusions. The guidance is inferred from the tool name and operation semantics rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_thread_getBRead-only
Ein Thread mit seinen Nachrichten, bodies read ON DEMAND from the mail store via each store_ref, never from SQLite. A message whose locator no longer resolves answers error message_moved while the rest of the thread still renders; a body whose re-read hash disagrees with the indexed sha256 comes back flagged stale:true.
| Name | Required | Description | Default |
|---|---|---|---|
| threadId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description adds crucial behavioral context: bodies are fetched on demand from store_ref, not SQLite, and handles stale/moved messages explicitly. This goes beyond the annotation, enriching the agent's understanding of potential failure modes and data freshness.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, ~2 sentences, well-structured, with key behavioral notes front-loaded. No unnecessary words. Slight brevity in parameter info doesn't hurt structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (on-demand fetching, error cases), the description covers behavior well, but lacks any parameter explanation and has no output schema. The agent can't know what the response contains beyond 'thread with messages'. Slightly incomplete for full usability.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for both parameters. It doesn't explain workspaceId or threadId meanings or formats at all; the agent only knows they're string IDs. This is a clear gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it retrieves a thread with its messages, specifying source details. It distinguishes from siblings like mail_threads_list, though not naming an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage for fetching a specific thread's messages, but doesn't explicitly state when to choose it over mail_threads_list or other mail tools. No exclusions or alternatives named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mail_threads_listARead-only
Die Korrespondenz-Warteschlange (P5, computed at query time): threads with their bucket derived from newest-message direction versus newest draft (needs_reply, drafted, done; recent and all as views), filterable by account or contact. savedViewId applies a G00 saved view; its stored filters merge underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| bucket | No | ||
| accountId | No | ||
| contactId | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true, the annotation already declares the tool is read-only. The description adds valuable behavioral context: buckets are computed at query time from newest-message direction versus newest draft, and savedViewId applies a saved view whose filters merge with explicitly named filters. This explains internal logic and merge behavior beyond what the annotation conveys. It does not address pagination or return format, but those are not required given the read-only nature and lack of output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence that front-loads the purpose and packs in details about bucket computation and saved view behavior. While concise, it uses cryptic identifiers like 'P5' and 'G00' without explanation, which may confuse an agent and reduce clarity. The structure is adequate but not fully accessible due to these unexplained codes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with five parameters and no output schema, the description covers the key aspects: what the tool returns (threads with buckets), how buckets are determined, filtering options (account/contact), and saved view behavior. It does not specify pagination, sorting, or return fields, but these are optional given the context. The description is sufficient for an agent to call the correct operation and understand the primary semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Given 0% schema description coverage, the description compensates reasonably. It explains the bucket parameter by enumerating possible values and how they are derived, and clarifies savedViewId's role in merging stored filters with explicit ones. It also mentions filterability by account or contact, giving meaning to accountId and contactId. It does not specify exact formats or whether parameters are required beyond workspaceId, but the core semantics are conveyed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists email threads with bucket classification based on message direction versus drafts. It names specific buckets (needs_reply, drafted, done) and views (recent, all), which gives concrete meaning to the operation. It does not explicitly contrast this with sibling tools like mail_thread_get or mail_drafts_list, but the purpose is clear enough to differentiate a thread-listing operation from a single-thread fetch or draft list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains that the tool can filter by account or contact and apply saved views, which implies usage for retrieving threads by these criteria. However, it does not provide explicit guidance on when to prefer this tool over alternatives such as mail_thread_get (for a single thread) or mail_drafts_list (for drafts). The mention of 'recent and all as views' hints at broader querying but lacks explicit exclusions or alternate tool routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mark_batch_paidA
Confirm the bank executed a generated batch: posts one outgoing payment per item through A14's recordPayment (debit 2000 Kreditoren / credit the debtor bank account), each settling its own vendor bill, atomically for the whole batch. Requires confirmation:true, because marking paid must reflect a genuine bank confirmation, never a side effect. Idempotent per batch: a re-confirm under the same key replays the same result and never double-pays. CONSEQUENCE: Confirms the bank paid the batch and settles every bill in it.
| Name | Required | Description | Default |
|---|---|---|---|
| batchId | Yes | ||
| bankTxnId | No | ||
| valueDate | Yes | ||
| workspaceId | Yes | ||
| confirmation | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and delivers richly. It discloses the accounting postings (debit/credit), atomicity across the whole batch, idempotency per batch (re-confirm replays the same result and never double-pays), and the consequence of settling every bill. This exceeds what annotations would typically convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a compact ~90-word paragraph that front-loads the main action and then adds mechanics, preconditions, idempotency, and consequence in logical order. The final CONSEQUENCE sentence is mildly redundant with the opening, but it can serve as a scanning summary for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex, consequential mutation (multi-payment posting, atomicity, idempotency, no output schema), the description covers purpose, mechanics, exact posting accounts, preconditions, and side effects. Minor gaps remain: no mention of return values, no description of failure behavior when confirmation is false, and the bankTxnId parameter's role is unexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explicitly explains the critical parameters: confirmation must be true, and the idempotencyKey semantics (re-confirm under same key replays result, no double-pay). The batchId is implied as the batch to confirm. workspaceId, valueDate, and bankTxnId are left to inference, but they are relatively self-explanatory names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Confirm the bank executed a generated batch') and goes on to describe exact mechanics: posts one outgoing payment per item via A14's recordPayment, debits 2000 Kreditoren, credits the debtor bank account, and settles each vendor bill. This is specific and clearly distinguishes it from siblings like payment_batch_transmit or discard_payment_batch.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states the precondition: requires confirmation:true, because marking paid must reflect genuine bank confirmation, never a side effect. It conveys the appropriate context (after the bank has actually executed the batch) but does not explicitly name alternatives or state when not to use it, so it falls just short of full exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
match_billA
Gleiche eine A17-Kreditorenrechnung 3-fach gegen die Bestellung und ihre Wareneingänge ab: compares, in CHF base Rappen, the value of goods RECEIVED-but-not-yet-billed at the POs own line prices against the bills base net, within a FIXED tolerance (2% or a CHF 1.00 floor, either passing). Within tolerance writes a matched po_match row, links the A17 bill_id, raises billed_qty (never above received_qty: no over-match) and, when the PO is fully billed, advances received -> closed. POSTS NOTHING: the bills expense/inventory + Vorsteuer posting is A17->A02 (P3), match_bill stores only the link. OVER tolerance BLOCKS auto-match: it PERSISTS a po_match row with status variance (the exception stays visible) and returns variance_exceeded. A user holding the post` capability may pass override:true to force an overridden match (records overridden_by). Refusals: zero receipts -> nothing_received; a bill from a different supplier -> invalid_reference; the same bill matched twice -> already_matched. match_bill is EXCLUDED from automation (a variance override is a human judgment).
| Name | Required | Description | Default |
|---|---|---|---|
| poId | Yes | ||
| actor | No | ||
| billId | Yes | ||
| override | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility for behavior. It discloses crucial traits: it writes a po_match row, links bill_id, raises billed_qty (never above received_qty), transitions PO to closed when fully billed, posts nothing financially, blocks on over-tolerance with variance status, supports override, lists specific refusal reasons, and notes automation exclusion. This is highly transparent for an operation with side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and informative, with the core purpose stated first and details following in a logical flow. It uses a long paragraph but every sentence adds necessary operational detail. Mixing German and English is slightly unconventional but does not impede understanding. It is appropriately sized given the complexity of the operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex matching tool with no output schema and no annotations, the description is remarkably complete. It covers success paths, failure modes (nothing_received, invalid_reference, already_matched), tolerance logic, side effects on POs, override behavior, and automation exclusion. All critical information an agent needs to invoke it correctly is present, including return value variance_exceeded.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explicitly explains the 'override' parameter (requiring 'post' capability and recording overridden_by) and mentions tolerance thresholds. However, it leaves workspaceId, poId, billId, actor, and idempotencyKey unaddressed despite their self-explanatory names. It adds meaningful context for one key parameter but doesn't fully bridge the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and object: 'Gleiche eine A17-Kreditorenrechnung 3-fach gegen die Bestellung und ihre Wareneingänge ab' (match an A17 creditor invoice three-way against the order and its goods receipts). It then states exactly what the tool does, including comparison logic, tolerance, and outcomes operação. It is distinct from sibling tools like match_three_way_create and match_three_way_override because it focuses on the A17 bill matching process, not generic three-way matching.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by giving detailed operational conditions: it is excluded from automation, override requires 'post' capability, and refusal scenarios. However, it does not explicitly mention alternatives or when to use this vs other matching tools, though the uniqueness of the action is clear. It does provide clear when-not-to-use via the 'EXCLUDED from automation' statement and the override conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
match_qr_paymentARead-only
Score a credit's facts against the open invoices, writing nothing: revalidates the reference (QRR mod-10 recursive / ISO 11649 mod-97, the same derivation A11 issues with), finds the invoice the reference names, and answers a confidence with the reason in words. 'high' is reserved for an exact reference AND an exact amount, where exact means the invoice's open amount or open + unpaid Mahngebühr (a Mahnung's QR part carries the invoice's own reference). Everything else is 'medium' (amount_short, amount_over, currency_differs) or 'none' (no_invoice, already_paid, ambiguous_reference, reference_invalid, no_reference). A mistyped check digit is reported as a typo and never ranks. Reads only.
| Name | Required | Description | Default |
|---|---|---|---|
| currency | No | ||
| reference | No | ||
| valueDate | No | ||
| amountMinor | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the readOnlyHint annotation by detailing the internal validation logic (QRR mod-10 and ISO 11649 mod-97), the exact conditions for 'high' confidence, and the handling of edge cases like typos in check digits. It also explains what happens in ambiguous cases ('ambiguous_reference', 'no_reference'). This adds significant behavioral context beyond just 'reads only'. However, it doesn't disclose potential side effects like performance (e.g., may be slow if many invoices) or whether it caches results, but for a read-only tool with such rich logical detail, a 4 is warranted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but very efficient: it front-loads the core action and read-only nature, then explains the algorithm, confidence levels, and edge cases in just a few sentences. Every sentence adds value, and there is no fluff. The structure flows logically from purpose to outcome classes, making it easy for an agent to parse quickly. It exemplifies concise yet comprehensive writing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters and no output schema, the description does an excellent job of explaining the main logic and expected outputs. It covers the key aspects: input validation, outcome classification, and edge cases. The only gap is that it doesn't describe the exact output structure (e.g., whether 'confidence' and 'reason' are separate fields), nor does it mention any side effects or performance considerations. Given the absence of an output schema, this incomplete output documentation is a minor but noticeable gap, preventing a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Given that schema description coverage is 0% and there are 5 parameters, the description must compensate. It does provide context for 'amountMinor' (used to check exact amount for 'high' confidence) and 'reference' (validated, used to find invoice), and indirectly mentions 'currency' through 'currency_differs' outcome. However, it does not explicitly explain the parameters 'workspaceId' (likely to scope the invoice search) and 'valueDate' (possibly relevant to date matching), leaving some ambiguity. Despite this, the description goes a long way to clarify the most important parameters, so a 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: scoring a credit's facts against open invoices to produce a confidence level and reason. It uses specific verbs ('Score', 'revalidates', 'finds', 'answers') and identifies the resource (open invoices) and the exact output type (confidence with reason). It also distinguishes itself from related tools like apply_qr_match and override_qr_match by emphasizing it writes nothing and is read-only, making its role in the QR matching workflow clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly explains when to use this tool: to evaluate a credit against invoices and get a confidence score. It implicitly distinguishes it from mutating actions like apply_qr_match or override_qr_match (it writes nothing), and the list of possible confidence outcomes ('high', 'medium', 'none') with their triggers provides guidance on when this evaluation is appropriate. However, it does not explicitly state when NOT to use it or name alternative tools for other scenarios (e.g., if you need to actually apply the match), so it misses a point for not providing explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
match_status_for_billARead-only
The read-only payment gate A18 / the mark-paid UI consult before settling a vendor bill. Returns the match status (unmatched | matched | partial | overridden | reversed), the active matchId when there is one, the total still-open quantity on the PO, and canPay: true only for an active matched / partial / overridden match. An unmatched bill returns canPay false (I04 default policy: a bill is not payable through the gate until matched). Writes nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| billId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description substantially exceeds the readOnlyHint annotation by enumerating the return fields (match status, matchId, open quantity, canPay) and explaining the conditional logic for canPay. It also explicitly says 'Writes nothing', reinforcing the read-only behavior. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each packing useful detail. The purpose and return fields are front-loaded, and the policy note is appended without fluff. Slightly dense but clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description fully outlines the return shape and the canPay logic. The I04 policy note adds policy context. For a simple two-parameter read-only status check, nothing critical is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explain the two parameters. It only implies 'billId' refers to a vendor bill and never mentions 'workspaceId'. With zero coverage the description should compensate, but it provides no parameter-level guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns match status for a vendor bill and is explicitly described as a read-only gate for the mark-paid UI. The verb 'returns' and the specific fields it returns distinguish it from mutating siblings like match_bill or match_three_way_create.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives the context 'consult before settling a vendor bill', which tells the agent when to call it. It does not explicitly name alternative tools or give when-not-to-use guidance, but the read-only framing strongly implies it is not for creating or modifying matches.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
match_three_way_createA
AGENT-FIRST. Persist a three-way match for a vendor bill (A17) against its purchase order (I01) and goods receipts (I02) when the evaluation is inside tolerance. Pass poId to pin the SAME candidate PO the evaluation was run against; otherwise the supplier oldest open PO is chosen, so an evaluate pinned to one PO and a create with no poId can land against a different PO. The engine RE-COMPUTES the evaluation under a row guard (the evaluation snapshot passed in is stored for audit but never trusted for the decision), writes an immutable match header + one line per PO line, increments billed_qty on every matched PO line exactly once, and marks the consumed I02 receipt lines. It POSTS NOTHING: the bill posting stays A17 -> A02. Refused when there is no candidate PO (no_candidate_po), nothing is open to bill (nothing_received), the bill has no CHF base figure yet (bill_not_convertible), the evaluation is out of tolerance (out_of_tolerance: use match_three_way_override), or the bill already has an active match (match_already_exists). allowPartial defaults to the workspace policy. Idempotent under idempotencyKey. Needs the purchasing.match capability.
| Name | Required | Description | Default |
|---|---|---|---|
| poId | No | ||
| billId | Yes | ||
| evaluation | No | ||
| workspaceId | Yes | ||
| allowPartial | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers: it discloses that the engine re-computes the evaluation under a row guard, treats the snapshot as audit-only, writes an immutable header, increments billed_qty exactly once, marks consumed receipt lines, posts nothing, is idempotent under idempotencyKey, and requires the purchasing.match capability. This goes far beyond what the schema or annotations could convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense yet efficient; every sentence carries new, essential information. It is front-loaded with the purpose and then covers nuances (PO pinning, re-evaluation, side effects, refusals, idempotency, permissions) without repetition or filler. The length is justified by the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the high complexity, no annotations, and no output schema, the description covers almost everything an agent needs: side effects, refusal reasons, idempotency, permission, and parameter semantics. The main gaps are the lack of a success return-value description and the unexplained workspaceId, which are minor against the depth already provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does for most parameters: poId (pinning behavior), evaluation (stored but never trusted), allowPartial (defaults to workspace policy), idempotencyKey (idempotency guarantee), and billId (clearly the vendor bill). However, workspaceId is not described, and allowPartial's true/false meaning is not fully explained, leaving small but real gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Persist'), a clear resource ('three-way match'), and the exact document set (vendor bill A17 against purchase order I01 and goods receipts I02). It also names the sibling that applies when the evaluation is out of tolerance, distinguishing itself from match_three_way_override and the other match_three_way_* siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when the tool is valid ('when the evaluation is inside tolerance'), names the alternative for out-of-tolerance cases ('use match_three_way_override'), and gives precise guidance on passing poId vs letting the system pick the oldest open PO. It also lists all refusal conditions with their error codes, leaving no ambiguity about when not to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
match_three_way_evaluateARead-only
PURE. Evaluate (never write) a three-way match for a vendor bill against its purchase order and goods receipts. Returns a structured MatchEvaluation: per PO line the ordered / received / already-billed / open-for-billing quantities, the PO unit price and extended value, the quantity variance and a line status; and at the header the summed expected value, the bill base-net that was billed, the aggregate price / value variance, the exact tolerances applied and an overall status suggestion (matched | partial | variance | nothing_received | no_candidate_po). Pass poId to pin the candidate; otherwise the supplier oldest open PO is chosen. Callable repeatedly; it writes no rows.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| poId | No | ||
| billId | Yes | ||
| receiptIds | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, and the description reinforces it with 'PURE', 'never write', and 'writes no rows' – consistent, not contradictory. It adds meaningful behavioral detail beyond the hint: the PO selection logic, repeatability, and the exact structure of the returned MatchEvaluation, including possible overall statuses. It does not disclose error conditions or permission requirements, but for a read-only evaluation tool the transparency is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded with 'PURE. Evaluate (never write)' and every sentence adds value. The lengthy enumeration of MatchEvaluation fields is justified because there is no output schema to carry that information. It could be better structured with line breaks, but it is not bloated or repetitive.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description does an excellent job specifying return values, per-line and header computations, tolerances, and overall status options. The PO auto-selection rule is also covered. However, the asOf and receiptIds parameters are not addressed, and workspaceId is silently required, leaving minor gaps for a tool with 5 parameters and no schema descriptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explicitly explains poId ('pin the candidate') and billId (the vendor bill being matched), but leaves workspaceId, receiptIds, and asOf entirely unexplained. The mention of 'goods receipts' hints at receiptIds but does not clarify whether it filters or overrides them. This partial compensation lands at a 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Evaluate (never write) a three-way match for a vendor bill against its purchase order and goods receipts' – a specific verb, resource, and scope. It distinguishes itself from siblings like match_three_way_create, match_three_way_override, and match_three_way_reverse by explicitly stating this is evaluation-only, and the detailed MatchEvaluation return type makes the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear direction on PO selection: 'Pass poId to pin the candidate; otherwise the supplier oldest open PO is chosen.' This tells the agent exactly how to control behavior. It also signals repeatability ('Callable repeatedly'), implying safe pre-decision assessment, though it does not name alternative sibling tools or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
match_three_way_exceptionsARead-only
The open exception list for prioritisation or human hand-off: posted, unmatched vendor bills whose live evaluation is variance or nothing_received, with the amount at risk (the value variance in Rappen), the supplier, and the age in days. Filter by supplierId and olderThanDays. Computed live (variance evaluations are never persisted).
| Name | Required | Description | Default |
|---|---|---|---|
| supplierId | No | ||
| workspaceId | Yes | ||
| olderThanDays | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint already set, the description adds meaningful behavior: results are computed live and variance evaluations are never persisted. This helps the agent understand that the list is dynamic and cannot be relied on as a stored snapshot. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, leading with the purpose before explaining filters and computation. Every sentence adds useful information with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema, the description names the key returned attributes (amount at risk in Rappen, supplier, age in days) and explains live computation. It leaves some gaps — workspaceId semantics, pagination/ordering, and exact olderThanDays comparison — but is adequate for a simple read-only list tool with annotations covering the safety profile.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explains that filters are supplierId and olderThanDays, and ties olderThanDays to the age-in-days concept. However, schema description coverage is 0%, and the required workspaceId parameter is not mentioned at all. It compensates partially but not fully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific resource — the open exception list of posted, unmatched vendor bills — and states the exact selection criteria (live evaluation is variance or nothing_received). It distinguishes itself from related siblings like match_three_way_list by emphasizing the exception/prioritisation angle and live computation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear context: the tool is for prioritisation or human hand-off of open exceptions, with optional filters by supplierId and olderThanDays. It does not explicitly name alternatives or say when not to use it, but the intended use case is explicit enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
match_three_way_getARead-only
Fetch one persisted three-way match: the header, its per-PO-line rows, and the stored evaluation snapshot (the exact numbers the GUI showed when the match was accepted).
| Name | Required | Description | Default |
|---|---|---|---|
| matchId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds useful behavior beyond the readOnlyHint annotation by emphasizing that the evaluation snapshot is the exact stored numbers shown when the match was accepted, indicating it is historical rather than recomputed. It doesn't cover errors, freshness, or authorization, but it goes beyond the annotation baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense, front-loaded sentence that names the operation, the resource, and the returned sections. Every clause adds meaning and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read operation with two required parameters and no output schema, the description adequately conveys what comes back: header, rows, and snapshot. It doesn't spell out parameter semantics or fail states, but the structure is simple enough that this does not feel incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of explaining matchId and workspaceId. It never mentions which identifier selects the match or how the workspace scopes the lookup; only the readable parameter names make it guessable. With zero schema descriptions, this is a real gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Uses a specific verb ('Fetch') and a well-defined resource ('one persisted three-way match'), then enumerates what is returned: header, per-PO-line rows, and stored evaluation snapshot. This clearly distinguishes it from siblings like match_three_way_list and match_three_way_evaluate without needing to open their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context: this is for retrieving a single already-accepted, persisted match and its historical evaluation snapshot, implying the agent should use list/evaluate siblings for those other purposes. It does not name alternatives explicitly or state when not to use it, but the singular and 'persisted' wording makes the scope obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
match_three_way_listARead-only
List persisted three-way matches, newest first, filtered by billId, poId, status (one value or a list) and a created-at from / to window.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| poId | No | ||
| billId | No | ||
| status | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds useful behavior beyond that: ordering by newest first, the ability to pass status as a single value or list, and the meaning of from/to as a created-at window. It does not mention pagination or response shape, but the read-only safety is covered by annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence that front-loads the verb and resource, then efficiently packs ordering and all filter semantics. Every clause earns its place with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main query intent and parameters, but for a list tool with no output schema it lacks guidance on pagination, result limits, or return format. Given the large sibling set and six parameters, a bit more completion would help an agent invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of explaining parameters. It maps billId, poId, status, and from/to to meaningful filters, clarifies that status accepts one value or a list, and defines from/to as a created-at window. It does not elaborate on workspaceId or date format, but the core filter semantics are present.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('List'), a precise resource ('persisted three-way matches'), and key behaviors ('newest first', filtered by billId, poId, status, and a created-at window). This clearly distinguishes it from sibling tools like match_three_way_get, match_three_way_create, and match_three_way_evaluate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when you need a filtered list of persisted three-way matches rather than a single match or a mutation. However, it does not explicitly name alternatives or state exclusions, such as using match_three_way_get for a single match or match_status_for_bill for status by bill.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
match_three_way_overrideA
Force a match on an out-of-tolerance evaluation with a MANDATORY reason (min 5 chars). Pass poId to pin the SAME candidate PO the evaluation was run against; otherwise the supplier oldest open PO is chosen. Writes a permanent match header with status overridden, the reason and the overriding actor, and still increments billed_qty and marks the receipt lines. An override without a reason is refused (reason_required). The override is permanent: it can only ever be undone by match_three_way_reverse, never silently turned back into a clean match. Refused with no_candidate_po / nothing_received / bill_not_convertible / match_already_exists exactly as create. Idempotent under idempotencyKey. Needs BOTH purchasing.match AND purchasing.match_override.
| Name | Required | Description | Default |
|---|---|---|---|
| poId | No | ||
| billId | Yes | ||
| reason | Yes | ||
| evaluation | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It details side effects ('increments billed_qty and marks the receipt lines'), permanence ('can only ever be undone by match_three_way_reverse'), mandatory reason, idempotency, required permissions, and error codes. This is exemplary for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the core action and reason requirement. It uses short, single-purpose sentences for side effects, errors, idempotency, and permissions, keeping each sentence non-redundant. It is slightly long, but the length is justified by the tool's complexity and the total absence of annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex mutation tool with no output schema, no annotations, and low schema coverage, the description provides a thorough operational picture: PO selection, side effects, permanence, error codes, idempotency, and permissions. The only notable gap is the unexplained optional 'evaluation' parameter, which the schema includes but the description never addresses.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It meaningfully explains poId (pin the same PO vs oldest open PO) and reason (mandatory, min 5 chars, refusal without it), but leaves evaluation, workspaceId, billId, and idempotencyKey without explicit semantic guidance. It adds value but does not fully cover all six parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Force a match on an out-of-tolerance evaluation.' It distinguishes from siblings by naming match_three_way_reverse as the only undo path and referencing 'exactly as create' for errors, so the tool's unique override role is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly implies usage when an evaluation is out-of-tolerance and explains PO pinning behavior ('Pass poId to pin the SAME candidate PO... otherwise the supplier oldest open PO is chosen'). It names match_three_way_reverse as the undo mechanism, but does not explicitly contrast with match_three_way_create for in-tolerance cases beyond the 'exactly as create' error reference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
match_three_way_reverseA
Reverse a matched / partial / overridden match when a material error is found. Writes a compensating record (status reversed) linked to the original, restores billed_qty on every affected PO line to its exact pre-match value, and un-marks the I02 receipt lines. It never mutates the original header or its lines (§H-AUDIT) beyond stamping the one-way reverse link. A reversed match cannot be re-activated: a fresh evaluate + create is required. A non-empty reason (min 5 chars) is mandatory (reason_required); a match that is already reversed or in a non-reversible state is refused (match_not_reversible). Idempotent under idempotencyKey. Needs the purchasing.match_override capability.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| matchId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so thoroughly. It spells out side effects (compensating record, restoring billed_qty, un-marking I02 lines), non-mutation guarantee, irreversibility, mandatory reason with min length, refusal conditions, idempotency, and required capability. This is exemplary disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense with every sentence earning its place: purpose, side effects, constraints, and capability. It is front-loaded with the core purpose and then systematically covers behavior. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the tool's behavior comprehensively: what it does, side effects, invariants, preconditions, error conditions, and idempotency. The only omission is the return value/response, but for a side-effect-heavy reversal action this is less critical and no output schema exists. The reference to '§H-AUDIT' is cryptic but does not block correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It adds meaning for reason (mandatory, min 5 chars), idempotencyKey (idempotent under it), and implies the matchId is the 'match' being reversed. It does not explicitly explain workspaceId or matchId, but these are standard identifiers and the domain context ('match') makes them self-evident.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Reverse a matched / partial / overridden match') and immediately names the trigger condition ('when a material error is found'). The operation is clearly distinct from sibling tools like match_three_way_create or match_three_way_override, so an agent can select it without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool (material error found) and states that a reversed match cannot be re-activated and requires a fresh evaluate + create. It does not explicitly name sibling tools as alternatives or list exclusions for when not to use this tool, but the context and constraints are sufficient for most decision-making.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_abandon_planC
Close a plan and discard its Testmandant. Belege for any step that reached live are retained (OR 958f), so abandoning never deletes the evidence of what was committed. After any commit this verb requires commit_migration.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| confirmed | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since there are no annotations, the description carries the full burden of behavioral disclosure. It does reveal two important traits: it discards the Testmandant, and it retains evidence of committed steps. It also mentions a permission requirement. However, it does not discuss reversibility, side effects on the plan state, or confirmation requirements beyond a vague mention of 'commit_migration'. More behavior could be described, making this adequate but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences and gets to the point quickly. The main action is stated up front. However, the reference 'OR 958f' is cryptic and unexplained, which may confuse agents. The phrase 'After any commit this verb requires commit_migration' is slightly awkward but still concise. Overall, it is efficient without being overly verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description is incomplete. It does not describe what 'close' entails beyond discarding the Testmandant, does not explain the required parameters, and does not mention the confirmation flag. The absence of any relationship to related migration tools (like 'migration_close_plan') makes it hard to place contextually. An agent might call this without understanding the full implications.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the four parameters (workspaceId, planId, idempotencyKey, confirmed). The description explains none of them. It does not define the purpose of 'confirmed' or the others. With zero schema descriptions and zero parameter explanation, this is a critical gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Close a plan and discard its Testmandant.' This is a specific verb+resource pair. However, it does not differentiate from siblings like 'migration_close_plan' or 'discard_testmandant', which could overlap in function. It fails to name alternatives, so it does not fully achieve the distinctiveness required for a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a usage constraint: 'After any commit this verb requires commit_migration.' It also reassures about data retention ('Belege for any step that reached live are retained'). However, it gives no explicit guidance on when to choose this tool over related siblings (e.g., when to use 'migration_close_plan' instead). The 'when-to-use' vs. alternatives is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_apply_map_templateA
Apply a Zuordnungsvorlage to a plan's maps and report the deltas: applied, unmatched (targets this workspace lacks), new (source accounts the template does not cover) and conflicts. A hand-set entry is never overwritten; the hand-set value wins.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| templateId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that the tool reports four delta categories and that hand-set entries win over template values – a key behavioral rule. However, it does not explain the idempotencyKey semantics or what happens on a conflict between template entries. Still, it adds substantial value beyond a simple 'apply' statement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences: the first front-loads the action and output; the second adds the critical hand-set precedence rule. There is no filler, and every word contributes to understanding the tool's purpose and behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description covers the core outcome (deltas) and a key behavioral rule. It does not explain the idempotencyKey purpose or detail what 'unmatched' and 'new' precisely mean, but it provides enough for an agent to know what to expect. Minor gaps remain, but it is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description does not explicitly explain the parameters. The context implies workspaceId, planId, and templateId are the identifiers for the workspace, plan, and template, but idempotencyKey is not mentioned at all. The description does not add meaning beyond the parameter names, and the idempotencyKey purpose is unclear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Apply a Zuordnungsvorlage to a plan's maps') and names the output ('report the deltas: applied, unmatched, new, conflicts'). It also clarifies a critical behavior (hand-set entries are never overwritten). This distinguishes it from siblings like migration_set_map (which sets a single mapping) and migration_save_map_template (which saves a template) – the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (when you have a mapping template to apply to a plan's maps) but does not explicitly mention alternatives or exclusion conditions. It does not say 'use this instead of X' or 'if you want to set a single mapping, use migration_set_map'. The context is clear enough for an agent to infer, but explicit guidance is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_check_stepA
Run the Eröffnungsprüfung for a step: every applicable control (trial balance balanced and per-account against the declared source, open AR/AP against their control accounts, each bank opening against the camt.053 OPBD, the VAT position at the Stichtag, row counts, document integrity, the export date against the Übernahmestichtag), persisted as an append-only snapshot whose hash a commit approval binds to. Idempotent: unchanged inputs return the stored snapshot. clean means every control passed or was waived; a failed or undeclared control refuses the commit.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| stepId | Yes | ||
| against | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses key behaviors: the check is append-only, idempotent (returns stored snapshot for unchanged inputs), and influences commit decisions. It also explains the meaning of 'clean' and the consequence of failure. This is more transparent than most tool descriptions and goes well beyond minimal required information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a dense, long sentence but it avoids fluff. It front-loads the purpose, then lists controls, and then states idempotency and commit implications. While it could be broken into clearer sentences, the structure is logical and every clause adds value. For a complex tool, this level of detail is justified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool that runs a complex set of checks)Skip it explains the scope of controls, persistence, idempotency, and the effect on commit. It does not describe the return value or error handling beyond commit refusal, but with no output schema and given the tool's role in a workflow, the description provides enough context for an agent to understand when to call it and what consequences it has. The missing part is explicit parameter definitions, which are partially covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies 'stepId' (the step being checked) and implicitly 'idempotencyKey' via the idempotency statement. It also references 'against' by describing checks 'against the declared source' and other control accounts. However, workspaceId and planId are not explicitly explained, though they are inferable as context identifiers for the migration plan. The mapping is partial but not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Run the Eröffnungsprüfung for a step') and enumerates exactly what controls are performed)Skip it clearly distinguishes itself from sibling tools like migration_get_check and migration_list_checks, which are about retrieving checks, by stating it runs the check and persists a snapshot.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage in the migration commit workflow: the check is persisted as a snapshot that a commit approval binds to, and a failed or undeclared control refuses the commit. This gives strong contextual timing. It does not explicitly name alternatives, but the description of idempotency and snapshot persistence clarifies that this tool is for executing the check rather than merely viewing existing ones.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_close_planA
Close a finished Datenübernahme, moving it from live to closed (the state the machine declared but nothing could reach). Legal only from live; refuses, naming the first blocker, while any step is non-terminal or any G11 control is failed. A close is the human judgment that the übernahme is finished, so it needs a confirmation and is denylisted from automation. Idempotent; writes no domain row and posts nothing (it stamps the plan only).
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| confirmed | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and excels: discloses idempotency, no domain row writes, no posts (only stamps the plan), refusal behavior with blocker naming, and the requirement for confirmation. These are critical behavioral traits an agent needs to know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with zero fluff. Purpose is front-loaded, and each sentence adds new, essential information. Well-structured and appropriately sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Comprehensive for a tool with no output schema. Covers purpose, preconditions, side effects, idempotency, confirmation requirement, and automation restriction. Nothing an agent needs to decide to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It clarifies the 'confirmed' parameter ('needs a confirmation') and the 'idempotencyKey' (idempotent), but doesn't explain 'workspaceId' or 'planId'. Partial compensation, but not full for all parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb 'Close' and resource 'Datenübernahme' (data migration), and clarifies the state transition from live to closed. It distinguishes itself from siblings like migration_abandon_plan by specifying the 'finished' condition and the human judgment requirement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit preconditions: legal only from live, refuses when any step is non-terminal or any G11 control is failed. Also states it requires confirmation and is denylisted from automation, which guides when a human should call it. However, it doesn't name alternative tools or explicitly say when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_commit_stepA
Commit a step into the live books, behind the six-condition gate (checked, no failed/not_asserted control, no unresolved conflict, an approval bound to the check hash for a money-path class, commit_migration held, and a G04 backup on record). A migrated document posts nothing: a money-path class establishes opening balances through A04's single entry. Idempotent: a double-commit posts once.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| stepId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses idempotency, that migrated documents post nothing, and how opening balances are set via A04. It does not cover error scenarios or reversibility, but the core behavior is well specified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact but dense with domain jargon (G04 backup, A04 single entry, money-path class). It front-loads the action but then buries the reader in specialized terms that may reduce clarity for an agent unfamiliar with the migration domain.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity and lack of output schema or annotations, the description covers the key aspects: what it does, the gate conditions, idempotency, and posting behavior. It omits response details and error handling, but these are not essential for deciding when to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must explain the parameters. It only hints at idempotencyKey via the idempotency note, but does not define workspaceId, planId, or stepId. The agent must infer their meaning from the tool name and context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Commit a step into the live books') and the resource (a migration step), and distinguishes itself from siblings like migration_check_step and migration_rollback_step. The gate conditions and idempotency behavior further clarify its role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description lists the six conditions that must hold before committing, which implicitly tells the agent when this tool is appropriate. It does not explicitly name alternatives (e.g., 'use migration_check_step first'), but the preconditions are concrete and act as guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_create_planA
Create a Datenübernahme plan for a source system as at an Übernahmestichtag. A FUTURE Stichtag is accepted: the plan is prepared ahead of its date (scope, map, trial load, check), the plan view carries cutoverPending, and only migration_commit_step and go_productive refuse cutover_in_future until the date arrives. The plan is a scoped, resumable, auditable object, not an import.
| Name | Required | Description | Default |
|---|---|---|---|
| localePack | No | ||
| cutoverDate | Yes | ||
| workspaceId | Yes | ||
| sourceSystem | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that future dates are accepted and how the plan is prepared ahead of the date, that the plan view carries cutoverPending, and that only migration_commit_step and go_productive refuse cutover_in_future until the date arrives. It also clarifies the plan is a scoped, resumable, auditable object. This is substantial behavioral disclosure, though it does not mention error handling, permissions, or what happens for past dates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, with the first stating the core purpose and the second adding behavioral detail and a clarifying exclusion. It is front-loaded, uses no filler, and each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description explains the plan creation and future-date behavior but does not describe the return value or plan structure in detail. It mentions the plan is scoped, resumable, and auditable, which gives context, but an agent might need to know what a successful response looks like. It is sufficient for understanding when and how to call it, but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 5 parameters and 0% description coverage. The description gives meaning to cutoverDate by explaining it can be a future Stichtag, and it refers to sourceSystem as 'source system'. However, it does not explain idempotencyKey, workspaceId, or localePack. It provides some semantic value but does not fully compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (create) and the resource (a Datenübernahme plan) with a specific scope (for a source system as at a cutover date). It distinguishes itself from import operations and other migration steps by noting that only migration_commit_step and go_productive refuse future cutover dates, which differentiates it from sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives context on when to use this tool: to create a plan, including for future dates. It explicitly says 'not an import', which implies that import tasks should use a different tool, but it does not list alternative tools by name or provide explicit 'when-not-to-use' scenarios beyond the import distinction. This is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_create_testmandantA
Provision the Testmandant for a plan: a disposable trial workspace (workspace.kind = sandbox) that behaves identically to a real one, composed from A00 create_workspace and linked to the plan so every migration_trial_load_step targets it. Idempotent per plan: a second call returns the existing Testmandant, never a second workspace; a plan already live refuses with plan_already_live. The user-facing word is Testmandant, never "Sandbox".
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it delivers: idempotency ('a second call returns the existing Testmandant, never a second workspace'), the error contract ('refuses with plan_already_live'), the side-effect chain (composed from create_workspace, linked to the plan so every migration_trial_load_step targets it), and even user-facing naming guidance ('The user-facing word is Testmandant, never "Sandbox"'). This is exceptional disclosure for a mutation tool with zero annotation coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, zero filler, with the core definition front-loaded. Sentence one defines the operation and its dependencies, sentence two covers idempotency and failure behavior, sentence three adds a small but actionable naming note. Every sentence earns its place and no information is repeated from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The behavioral contract (idempotency, error state, linking, composition) is thoroughly covered, which matters given no annotations exist. However, there is no output schema and the description never states what the tool returns (the Testmandant workspace? a status object?), and the unexplained workspaceId parameter leaves a real gap for a tool with three required parameters. Strong on behavior, incomplete on return value and parameter roles.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It gives planId clear meaning ('for a plan', 'linked to the plan') and conceptually explains idempotency, but workspaceId is never explained — it's unclear why a provisioning tool that creates a workspace requires a workspaceId — and idempotencyKey's relationship to the stated 'idempotent per plan' behavior is left ambiguous. The description adds substantial context overall but leaves two of three required parameters under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('Provision the Testmandant for a plan') and then defines the resource precisely: a disposable trial workspace (workspace.kind = sandbox) that behaves like a real one. It distinguishes itself from nearby siblings by explicitly naming create_workspace as its composition source and migration_trial_load_step as the dependent that targets it. An agent can tell exactly what this tool does and how it relates to the migration_* family.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use it: provisioning the trial workspace for a migration plan, and it implicitly says don't call create_workspace directly since this tool is 'composed from A00 create_workspace.' It also states a when-not-to-use condition ('a plan already live refuses with plan_already_live') and that it must precede migration_trial_load_step. It stops short of an explicit alternatives framing (e.g., naming migration_get_testmandant for retrieval or discard_testmandant for removal), so it doesn't earn a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_declare_control_totalA
Declare an expectation for one Eröffnungsprüfung control: the figure the old system said (a per-account trial-balance total, the open AR or AP total, the VAT position at the Übernahmestichtag), in integer Rappen against a scope (an account number, an IBAN, or workspace). A control nobody declared reports "nicht geprüft" and never green; declaring reopens the control until the next check.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | ||
| scope | Yes | ||
| planId | Yes | ||
| stepId | No | ||
| workspaceId | Yes | ||
| declaredMinor | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that declaring reopens the control and that undeclared controls are marked 'nicht geprüft'. However, it does not mention idempotency behavior, whether repeated declarations overwrite, error conditions, or permission requirements. While it adds useful context, it is not exhaustive for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused paragraph that front-loads the core purpose, includes relevant examples, and is free of redundancy. Every sentence adds value, and it is appropriately concise for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 7 parameters, 0% schema coverage, no output schema, and no annotations, the description is not complete. It explains the core concept and some behavior but leaves several parameters unexplained and omits return values, error handling, and relationships to sibling migration tools. An agent would likely need to guess at parameter formats and side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain all parameters. It clarifies that 'declaredMinor' is the figure in integer Rappen and that 'scope' is an account number, IBAN, or workspace, but it does not explain 'kind', 'planId', 'stepId', 'workspaceId', or 'idempotencyKey'. With 7 parameters and no schema descriptions, this is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: declaring an expectation for a control in the Eröffnungsprüfung, with specific examples of what the figure represents (trial balance total, AR/AP total, VAT position) and the unit (integer Rappen). It differentiates from sibling tools like migration_waive_control or migration_check_step by focusing on declaration, making it unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use the tool (to declare an expected control value) and the consequence of not declaring (reports 'nicht geprüft' and never green), plus that declaring reopens the control. It does not explicitly name alternatives, but the purpose is clear enough to infer that this is for setting expectations rather than waiving or checking. Slight lack of explicit when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_diff_testmandant_to_liveARead-only
What differs between the Testmandant and an existing live workspace of the same UID, per data class: the row counts on each side, so the choice between going productive and running against the live workspace is informed. No live workspace of this UID returns {live:null}. This read touches TWO tenants and demands membership of EACH (forbidden otherwise): it is the migration family's one deliberate two-workspace read. It never merges.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses that the read touches two tenants and requires membership in each, that a missing live workspace returns {live:null}, and that the operation never merges anything. These are meaningful behavioral details beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
All four sentences earn their place: main output, null edge case, cross-tenant authorization behavior, and non-mutation guarantee. The most important semantic is front-loaded, and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The behavior is richly described, but the tool cannot be correctly invoked from this description alone because neither required parameter is semantically defined. There is also no output schema, and the return shape is only described conceptually as row counts per data class.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it never explains the two required parameters, workspaceId and planId. It references 'same UID' and 'Testmandant' without mapping them to the actual parameters, and planId is entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states what the tool computes: differences in row counts per data class between the Testmandant and a live workspace of the same UID. It also differentiates this tool from other migration-family tools by calling it 'the migration family's one deliberate two-workspace read.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: it informs the choice between going productive and running against the live workspace. It also implicitly excludes other migration tools by emphasizing this is the only deliberate two-workspace read, but it does not explicitly name alternatives or list when-not-to-use conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_discover_sourceARead-only
Classify uploaded source files for a Datenübernahme: per file the detected adapter, the data classes it can produce, a row count, a header sample, a confidence and the file's own as-of date (asAt) where the format carries one. Pass override to FORCE a format when auto-detection got it wrong: override.adapter pins a registered adapter id, override.encoding re-decodes (utf-8, latin1, windows-1252) and override.delimiter re-splits (comma, semicolon, tab). Reads blobs back through E00 and writes zero rows to any ledger.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | No | ||
| fileIds | Yes | ||
| override | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true, and the description adds critical behavioral information not covered by annotations: it reads blobs through E00 and writes zero rows to any ledger. This reassures that the operation is safe for the system state, complementing the read-only annotation. There is no contradiction; the description explains the side-effect-free nature of the classification process.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, with a single paragraph that front-loads the core purpose and outputs, then explains the override mechanism and ends with side-effect notes. Every sentence adds value, but the structure could be improved by separating the override details into a list. Overall, it's efficient and clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (nested objects, multiple files, overrides) and the absence of an output schema, the description covers most essential aspects: what it does, what it returns per file, how to correct errors, and side effects. The main gap is the lack of explanation for 'planId' and the exact return format beyond mentions. But it's reasonably complete for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. The description explains 'override' in detail (pinning adapter, re-decoding with specific encodings, re-splitting with delimiters) and mentions 'fileIds' and 'workspaceId' indirectly ('uploaded source files'). However, 'planId' is not mentioned, nor are the exact formats of the parameters. The description adds value for 'override' but leaves 'planId' unexplained, which is a gap given low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: classifying uploaded source files for a Datenübernahme, with specific outputs (detected adapter, data classes, row count, etc.). The verb 'classify' and the resource 'uploaded source files' are specific and distinguish it from other migration tools like migration_rollback_step or migration_readiness. Although sibling tools are numerous, this one's purpose is unique among the migration family.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use the tool: for classifying source files during a Datenübernahme, and it describes the override parameter as a corrective action when auto-detection fails. It doesn't explicitly state when not to use it or name alternatives, but the migration sibling tools like migration_list_source_adapters or migration_preview_step suggest alternative purposes. The context is clear enough for an agent to infer appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_export_checkARead-only
Export the Prüfbericht for one check: every control with its declared and computed figures, its status, every waiver and its reason, the source files with their sha256 and as-of date, and the check hash, as a locale-neutral artifact (raw integer Rappen, ISO dates) a Treuhänder can file. It joins the A25 filing export set rather than becoming a second export mechanism.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | ||
| checkId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the exact output composition: controls with declared/computed figures, status, waivers with reasons, source files with sha256 and as-of date, and check hash. It also specifies the locale-neutral format (raw integer Rappen, ISO dates). This goes well beyond the readOnlyHint annotation, explaining the artifact's nature and purpose for filing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, information-dense sentence that front-loads the core purpose ('Export the Prüfbericht for one check') before listing details. It is efficient, though slightly long; every phrase adds value by specifying content and format.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description fully explains what the output contains and its purpose, which is helpful given no output schema. However, it omits parameter semantics, potential errors, prerequisites (e.g., check must exist), and any rate limits or size constraints. The complexity is moderate, but the missing parameter details leave an agent uncertain about how to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate for parameter meanings, but it does not mention any of the three parameters (format, checkId, workspaceId). checkId and workspaceId are somewhat inferable, but 'format' is entirely unexplained. The description offers no guidance on parameter values or constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool exports a specific report (Prüfbericht) for a single check, enumerating its contents (controls, waivers, source files, check hash). It distinguishes itself from sibling tools like migration_get_check (which presumably retrieves check data) and other export tools by framing it as a filing artifact that joins the A25 export set.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: when a locale-neutral, filable export of a check's audit report is needed for A25 filing. It explicitly says it 'joins the A25 filing export set rather than becoming a second export mechanism,' providing contextual differentiation from other exports. However, it doesn't mention when not to use it or alternative tools like migration_get_check for simple retrieval.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_get_checkARead-only
Read one persisted Eröffnungsprüfung: every control with its declared and computed figures in Rappen, its status (passed, failed, not_asserted, not_computable, waived), and every waiver with its recorded reason. The snapshot is append-only: this is the record an approval was bound to.
| Name | Required | Description | Default |
|---|---|---|---|
| checkId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds significant behavioral context beyond the readOnlyHint annotation: it specifies that the snapshot is append-only and that it represents the record an approval was bound to. This tells the agent the data is immutable and historical, which is not evident from the annotation alone. It does not contradict the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, but each packs essential information. The first sentence lists the expected contents, and the second adds the append-only nature and purpose. It is front-loaded with the core action and resource, making it easy to scan. No redundant words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (it returns a detailed snapshot with multiple fields) and the lack of an output schema, the description provides a good overview of what to expect. It mentions all key elements: controls, figures, statuses, waivers, and reasons. However, it does not mention potential pagination, error conditions, or why a check might not be found, which could be relevant. But for a read operation with a clear purpose, it is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and there are only two parameters: workspaceId and checkId. The description mentions the check but not the workspace, so it does not add semantic detail for the parameters. However, the parameter names are self-explanatory: workspaceId likely identifies the workspace, and checkId identifies the check. The description's mention of 'one persisted Eröffnungsprüfung' implies checkId is the primary identifier, but it does not add beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads a persisted Eröffnungsprüfung and lists the exact contents: controls with declared/computed figures in Rappen, statuses, and waivers with reasons. It distinguishes itself from sibling tools like migration_check_step and migration_list_checks by focusing on reading a single check with full detail. The verb 'Read' and the specific resource make the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for retrieving a persisted check, likely after approval or for audit. It mentions 'append-only' and 'the record an approval was bound to,' suggesting it is used for historical reference. However, it does not explicitly state when to use this vs. migration_check_step or migration_list_checks, nor any conditions for when not to use it. The context is clear enough for an agent to infer typical usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_get_extraction_guideARead-only
Read one extraction guide for a source system: every item with what to export, where it lives, the expected formats, the data classes it feeds, quirks and its tactic-ladder rung (1 native export, 2 report-based, 3/4 browser companion (gated), 5 the Datenherausgabe letter). An unknown source system returns the generic guide with fellBack:true; pass strict:true to require an exact match. Includes the deletion clock and the letter template. Describes the software, not any workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| strict | No | ||
| sourceSystem | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already signals a safe read operation, and the description adds meaningful behavioral context: fallback to a generic guide with fellBack:true, strict matching behavior, and the inclusion of the deletion clock and letter template. This goes beyond the annotation by explaining edge-case behavior and return semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: it front-loads the core purpose, then details contents, fallback behavior, strict mode, and scope clarification. Every sentence adds information; there is no filler or repetition of schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with no output schema, the description covers the key behaviors an agent needs: what the guide contains, fallback behavior, strict matching, and scope. It doesn't describe the exact return structure, but the description's enumeration of contents (formats, data classes, quirks, tactic-ladder rung, deletion clock, letter template) gives sufficient context for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of explaining parameters. It explains sourceSystem (identifies the source system, unknown values fall back to generic guide) and strict (requires exact match). This is meaningful semantic context beyond the bare schema types. It doesn't explicitly name the parameters, but the behavior described maps clearly to them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read') and resource ('one extraction guide for a source system'), and enumerates the guide's contents (what to export, where it lives, expected formats, data classes, quirks, tactic-ladder rung). It also distinguishes itself from the sibling migration_list_extraction_guides by focusing on reading a single guide rather than listing guides.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly explains the fallback behavior for unknown source systems, the strict parameter to require an exact match, and what the guide includes (deletion clock, letter template). It also clarifies the tool describes the software, not any workspace, which helps an agent decide when to use it. While it doesn't name an alternative tool explicitly, the sibling list includes migration_list_extraction_guides, and the description's focus on 'one extraction guide' versus listing guides provides clear usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_get_manifestARead-only
Read a plan export manifest: its items with status, counts and evidence, the completeness (open count over the denominator, not_used excluded) and the deletion clock as a date with the days remaining. Accepts a saved view (savedViewId) over the checklist items ("Offene Exporte", "Blockiert"); an explicit status wins over the stored one. Completeness and the deadline are computed over the whole manifest, never the filtered view. A plan with no manifest returns manifest:null so the surface can offer to create one.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| status | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is transparent about the read-only nature (consistent with readOnlyHint=true) and discloses key behaviors: the effect of filtering on computed metrics, the precedence of explicit status, and the null return when no manifest exists. This adds significant context beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and information-dense, with each sentence serving a purpose: purpose statement, parameter behavior, computation scope, and edge case handling. It is well-structured and front-loaded with the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential aspects: what it returns, how parameters affect output, and a notable edge case. It does not explicitly mention that planId and workspaceId are required (though the schema does), and it omits potential error scenarios beyond the null manifest. Given the absence of an output schema, it adequately informs an agent of the tool's behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description compensates by explaining savedViewId and status in detail, including their roles and precedence. It does not explicitly describe planId and workspaceId, but these are standard identifiers and likely self-explanatory in context. Overall, it adds meaningful value over the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Read a plan export manifest' and elaborates on the specific data returned (items, status, counts, completeness, deletion clock). It distinguishes this from siblings like migration_get_plan or migration_set_manifest by focusing on the manifest read operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides concrete usage guidance on how savedViewId and status interact ('an explicit status wins over the stored one') and clarifies that completeness and deadline are computed over the whole manifest, not the filtered view. It does not explicitly contrast with alternative tools, but the parameter behavior is well explained.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_get_mapARead-only
Read a plan's map and its completeness: blocking unmapped accounts (non-zero balance), ignorable ones (zero balance), many-to-one collapses and template conflicts. complete:true is what lets the Datenübernahme step advance.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | ||
| planId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces that with 'Read' rather than contradicting it. It adds meaningful behavioral detail beyond the annotation by describing what the result contains and how 'complete:true' gates the next migration step.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the first sentence states the action and resource, and the second adds the key gating consequence. The only minor issue is the untranslated 'Datenübernahme' jargon, which may obscure the meaning for an agent not familiar with the migration domain.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With three required parameters, 0% schema description coverage, and no output schema, the description is not complete enough for an agent to reliably call the tool. It explains what the map contains, but leaves 'kind' and workspace/plan identifier semantics undefined and does not describe the response shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the bare parameters workspaceId, planId, and kind. It never explains kind, workspaceId, or how to identify the plan beyond the phrase 'a plan's map', leaving required parameter semantics essentially undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read') and resource ('a plan's map'), and enumerates the meaningful output categories: blocking unmapped accounts, ignorable accounts, many-to-one collapses, template conflicts, and the completeness flag. It does not explicitly contrast itself with siblings like migration_get_plan or migration_get_check, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: the tool is for checking a plan's map and its completeness, and the 'complete:true' flag is what lets the Datenübernahme step advance. This implies when to call it, but it does not state exclusions or name alternatives such as migration_suggest_map or migration_set_map.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_get_planARead-only
Read a plan, every step with its state, the six-phase journey and nextAction, the resume answer and the plan surface's single primary action. Accepts a saved view (savedViewId) over the step list; an explicit dataClass/state filter wins over the stored one. Resume is a read and never a re-run of prior steps.
| Name | Required | Description | Default |
|---|---|---|---|
| state | No | ||
| planId | Yes | ||
| dataClass | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations include readOnlyHint=true, and the description reinforces this by explicitly stating 'Resume is a read and never a re-run of prior steps.' This adds critical context beyond the annotation, clarifying a potential misunderstanding about the 'resume' concept. The description also discloses the filter-priority behavior (explicit filter wins over saved view), which is behavioral context not in the schema. No contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph covering all key points: what is read, the savedViewId behavior, filter override, and the resume caveat. It is front-loaded with the main purpose. Slightly verbose in listing the six-phase journey and nextAction, but each item adds clarity. No unnecessary fluff, though it could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool reads a complex plan with many components, the description covers the main return elements and key behaviors (savedViewId, filter override, resume semantics). There is no output schema, so the description carries full weight for return values; it lists the key items. It could mention whether the response includes step-level details like status codes or error handling, but for a read tool, this is adequate. It also doesn't discuss pagination or large plan handling, but that's minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. The description clarifies the semantics of savedViewId (filter over the step list) and mentions dataClass/state as filters, but it doesn't detail the exact format or meaning of each parameter. For required workspaceId and planId, it's clear they identify the plan. However, it omits how state and dataClass values should be formatted or what they represent beyond being filters. With five parameters and zero schema help, this is a moderate gap, but the description does add value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads a migration plan and lists the specific components it returns (steps, six-phase journey, nextAction, resume answer, primary action). It distinguishes itself from sibling tools like migration_list_plans and migration_check_step by focusing on a single plan's full detail. The verb 'Read a plan' is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is the read-only tool for a single plan, and contrasts with migration_list_plans (listing) and migration_rollback_step (acting). It explains the savedViewId behavior and the filter override rule, giving clear context. However, it doesn't explicitly state when to use an alternative like migration_get_check or migration_readiness, though sibling names suggest those are for checks and readiness.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_get_testmandantARead-only
Read the Testmandant for a plan: its workspace id, its kind (demo|sandbox|live), the steps trial-loaded into it, and the plan's last Eröffnungsprüfung id. A plan with no Testmandant returns {none:true}.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation is present, and the description adds meaningful behavior beyond it: it states the return fields, the allowed kind values, and the {none:true} sentinel when no Testmandant exists. This gives the agent concrete expectations without contradicting the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler: the read action is front-loaded, the return contents are compactly enumerated, and the edge case is stated at the end. Every part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with a simple two-string schema, the description covers the return shape and the no-testmandant case well. The main gap is the lack of explicit parameter meaning, which keeps it just short of fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not define either input parameter. 'For a plan' hints at planId, but workspaceId is never explained, leaving required-string parameter semantics mostly to guessing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb ('Read'), identifies the resource ('Testmandant for a plan'), and enumerates the exact fields returned, including the edge case. This clearly distinguishes it from sibling operations like migration_create_testmandant, migration_diff_testmandant_to_live, and discard_testmandant.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It is clear this is the read-only retrieval tool for a plan's Testmandant, and the context is fairly obvious. However, there is no explicit guidance about when to choose it over related siblings or when not to use it; usage is implied rather than directly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_list_checksARead-only
List a plan's Eröffnungsprüfungen, newest first, optionally for one step or one run kind (testmandant or live). Accepts a saved view (savedViewId) over the check history; an explicit stepId/against filter wins over the stored one.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| stepId | No | ||
| against | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint: true, and the description adds ordering (newest first) and filter precedence (explicit stepId/against wins over saved view). These are useful but modest additions. It does not disclose pagination, response format, or any edge cases, though for a read-only list tool this is acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no redundancy. The core action is front-loaded, followed by essential filtering and precedence rules. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main aspects: what is listed, ordering, optional filters, saved view interaction, and precedence. Given it is a read-only list tool with no output schema, it is sufficiently complete for an agent to invoke correctly. Minor gaps like pagination behavior are not critical.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the full burden. It clarifies stepId, against (run kind: testmandant or live), and savedViewId, including precedence. workspaceId and planId are left to inference but are standard and obvious from the tool name. The description compensates well for the missing schema details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it lists a plan's Eröffnungsprüfungen (opening checks) newest first, with optional filters. The verb+resource is specific and unambiguous, though it doesn't explicitly differentiate from sibling tools like migration_get_check (which fetches a single check) or migration_readiness. The German term may be opaque but is domain-specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains optional filtering by step or run kind, and the saved view precedence, giving context for when to use these options. However, it does not explicitly state when not to use this tool or name alternatives, leaving the agent to infer from the tool name and sibling context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_list_extraction_guidesARead-only
List the extraction guides TILL ships (one per source system, generic and bexio first), each with its label, item count and whether a browser companion has cleared its gates (hasCompanion). Describes the software, not any workspace.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds useful behavioral context: it lists per source system, orders 'generic and bexio first', and includes a boolean indicating whether a browser companion has cleared its gates. It does not, however, describe the full return format or any potential edge cases (e.g., what happens if no guides exist), but this is a minor gap given the simple read-only nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no wasted words. It front-loads the verb and resource, immediately provides the ordering rule, then lists the returned fieldsamerican. The final scoping sentence ('Describes the software, not any workspace') is a valuable clarification that prevents confusion with workspace-scoped list tools. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with no parameters and no output schema, the description covers the essentials: what is listed, the ordering, the fields returned, and the scope. It does not mention pagination or a maximum result count, but given the nature of 'extraction guides' (likely a limited set), this is not a critical gap. Overall, the information provided is sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters)Skip, so the schema already fully documents all inputs. With 100% schema coverage and no parameters, the description has no obligation to explain parameters. The baseline for 0 parameters is 4, and the description adds no confusing parameter-related information, so a 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('List'), the resource ('extraction guides'), and key output details (label, item count, hasCompanion). It also scopes itself ('Describes the software, not any workspace'). While it does not explicitly name sibling alternatives, the level of specificity is high enough for an agent to distinguish it from migration_get_extraction_guide or other migration tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance is given about when to use this tool versus alternatives. There is no mention of prerequisites, context (e.g., before a migration), or exclusions. The only implicit cue is the name and the mention of 'extraction guides' in the migration domain, which suggests a read-only listing operation but leaves the agent to infer the appropriate triggering conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_list_locale_packsARead-only
List the registered locale packs (target chart, tax-code set, parse conventions, statutory anchors). Switzerland (ch) always ships. Describes the software, not any workspace.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
readOnlyHint=true already establishes that this is a safe read, and the description adds context beyond the annotation: it guarantees Switzerland always appears in results, clarifies that no workspace context is involved, and reveals the conceptual contents of a locale pack. This usefully shapes agent expectations about the response despite the absence of an output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with zero filler. The purpose is front-loaded, followed by a factual guarantee (Switzerland always ships) and a scope boundary, each sentence earning its place without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only listing with no output schema, the description covers purpose, result composition, guaranteed entries, and scope. The only notable omission is the exact return format, but the enumerated locale-pack components largely compensate by implying what the response will contain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4 and there is nothing for the description to document. The empty schema already fully covers parameter semantics, and the description correctly implies a parameterless call by stating only what the tool does.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description pairs a specific verb and resource ('List the registered locale packs') and defines the resource by enumerating its components (target chart, tax-code set, parse conventions, statutory anchors). The closing scope clause ('Describes the software, not any workspace') helps differentiate it from workspace-scoped migration siblings, though no sibling is explicitly named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'not any workspace' clause gives an implicit exclusion, signaling this tool reports software-level facts rather than workspace state. However, there is no explicit when-to-use guidance, no named alternative, and no context such as 'call before starting a migration' to help an agent choose it over related migration listing tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_list_map_templatesARead-only
List this operator's Zuordnungsvorlagen, optionally filtered by source system or map kind. Accepts a saved view (savedViewId) over the template list; an explicit filter wins over the stored one.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| sourceSystem | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description is consistent with that. It adds useful behavioral detail beyond annotations by explaining that a saved view can be applied and that an explicit filter takes precedence over the stored view.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two focused sentences with no wasted words. The primary action and filters are front-loaded, followed by the saved-view precedence caveat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool, the description covers the core operation, filters, and saved-view behavior, but gaps remain: workspaceId is not explained, kind values are unspecified, and there is no mention of result shape or pagination. Since the schema has no property descriptions, the description needed to compensate more than it does.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaning to sourceSystem, kind, and savedViewId by mapping them to filters and a saved view, which matters given 0% schema coverage. However, workspaceId, the only required parameter, is not explained beyond its name, and allowed values for kind are not described.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear operation, listing the operator's mapping templates, with optional filters by source system or map kind. The resource is specific enough to distinguish it from sibling tools like migration_apply_map_template and migration_save_map_template, though it does not explicitly name a sibling. Use of the German term 'Zuordnungsvorlagen' adds mild ambiguity but the intent is still recognizable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is used to list templates and that saved views can refine the listing, which gives some usage context. It does not explicitly state when to prefer this over related migration mapping tools or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_list_plansARead-only
List this workspace's Datenübernahme plans, optionally by status. Accepts a saved view (savedViewId) over the roster; an explicit status wins over the stored one. Workspace-scoped like every other read: the cross-client Treuhänder roster composes N of these over A23's memberships and is NOT a read that reaches across the tenant fence.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces this with 'List' and 'read'. It adds meaningful context beyond the annotation by clarifying the workspace scope and the saved-view precedence behavior. It does not discuss rate limits, auth, or side effects, but the readOnlyHint covers safety and the added scope details are valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is reasonably compact with a clear front-loaded purpose, but the third sentence introduces domain-specific jargon (Datenübernahme, Treuhänder, A23, tenant fence) that is dense and potentially confusing to an agent without migration-specific context. It is not fully concise and may distract from the core usage guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and zero schema description coverage, the description is the primary source. It explains the key parameters and workspace scope, but it does not mention allowed status values, pagination, return format, or what constitutes a saved view. These are notable gaps for a list operation, especially without an output schema to compensate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden for parameter semantics. It explains status (optional filter), savedViewId (saved view over the roster), and workspaceId (workspace scope, implied by 'this workspace's'). It also adds a precedence rule: explicit status wins over the saved view's stored status. This goes beyond the bare schema and provides actionable guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists this workspace's migration plans, optionally filtered by status. It also distinguishes itself from cross-tenant aggregations by explicitly stating it is workspace-scoped and not a read that crosses the tenant fence, which separates it from potential sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this is a per-workspace read and not a cross-tenant operation. It implies when to use it (listing plans in one workspace) and when not (when you need a composite cross-client view), but it does not explicitly name an alternative tool. The guidance is clear but lacks an explicit exclusionary pointer to a specific sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_list_source_adaptersARead-only
List the source formats TILL can read for a Datenübernahme, each with the published documentation it was written from (cleanRoomSource). Describes the software, not any workspace.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes that this is a safe read operation. The description adds context about the output contents ('each with the published documentation it was written from') and clarifies that it pertains to software, not a workspace. This goes beyond the annotation but does not disclose other behavioral aspects like return format, pagination, or error conditions. Given the simple nature of the tool, the added context is sufficient for a 3.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: two sentences with no filler. The core action and resource are front-loaded, and the second sentence adds a clarifying scope note. Every word earns its place, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only list operation with no output schema, the description adequately conveys what the tool returns (source formats with their documentation) and its non-workspace scope. It does not describe pagination or limits, but these are unlikely to be critical for a simple enumeration. The information provided is sufficient for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description is not required to explain any. The schema is empty and there is nothing to add beyond the tool's purpose. The baseline for zero-parameter tools is 4, and the description meets that expectation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List') and a clear resource ('source formats TILL can read for a Datenübernahme'). It also adds a clarifying detail about the scope ('Describes the software, not any workspace'), which helps distinguish it from workspace-specific tools. However, it does not explicitly name sibling tools, so it misses the highest differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. It does not mention prerequisites, typical use cases, or conditions that would select this over related migration tools like migration_discover_source or migration_create_plan. The only implied usage is that it is a listing operation, but no explicit context is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_preview_stepARead-only
Dry-run a step: classify every source row willCreate, willSkip (an exact match on the class match key) or willConflict (a match with differing fields), with a sample and any row errors. Writes NOTHING anywhere; an unresolved conflict blocks the commit.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| stepId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces this with 'Writes NOTHING anywhere', which is consistent. It also adds the behavioral consequence that an unresolved conflict blocks the commit, which is not in the annotations and is valuable context. It does not describe return format or edge cases, but for a read-only dry-run with annotations covering safety, this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. It front-loads the core action ('Dry-run a step'), explains the classifications, and then states the key safety guarantee and consequence. Every sentence contributes value, and it is well-structured for quick parsing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a dry-run tool with no output schema, the description covers the essential aspects: what it does, what it returns (sample and row errors), and the impact on commit. It does not describe preconditions (e.g., step must exist) or how to interpret the classifications, but given the tool's simplicity and the presence of readOnlyHint, it is fairly complete. A missing element is any note about the need to run this before commit, which is implied by the conflict-blocking statement.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has three required parameters (workspaceId, planId, stepId) with no descriptions (coverage 0%). The tool description does not explain the meaning of these parameters or how they relate to each other, nor does it provide context on how to obtain them. While the names are intuitive, the lack of any elaboration means the agent must rely on naming conventions alone, which is insufficient for 0% coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: a dry-run that classifies each source row into willCreate, willSkip, or willConflict, and returns a sample and row errors. It also explicitly notes that it writes nothing, distinguishing it from commit-style operations. The purpose is unambiguous and specific, and it stands apart from sibling migration tools like migration_commit_step or migration_trial_load_step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: before committing a migration step, as a preview to identify conflicts. It explicitly mentions that an unresolved conflict blocks the commit, signaling the need to run this and resolve conflicts first. It does not explicitly name alternatives, but the context is clear and the safety property (writes nothing) makes it the appropriate pre-commit check.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_readinessARead-only
Read the go-live checklist: every step with its state, the backup precondition, and for each open item the actor who must move it (du, ein Agent, das System). Ready means every included step is verified and no control is failed or not_asserted.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation indicates a safe read operation; the description adds value by specifying exactly what the tool returns (step states, backup precondition, actor assignment) and defines the 'ready' condition. This goes beyond the annotation without contradicting it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core action, and efficiently conveys the checklist's content and the readiness criterion. No wasted words; every sentence contributes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool, it explains the output content and the readiness definition, but it omits any explanation of the parameters and does not mention potential edge cases (e.g., empty checklist, missing plan). The lack of parameter documentation is a notable gap, so completeness is only moderate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has zero description coverage for the two required parameters (workspaceId, planId). The description does not explain what these parameters represent or their format, forcing the agent to infer from context. Given low coverage, the description fails to compensate, leaving a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Read' and the resource 'go-live checklist', and elaborates on its contents (steps, states, backup precondition, actors). It is distinct from sibling migration tools which perform actions (rollback, abandon, check) rather than a read-only overview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for checking readiness but does not explicitly state when to use this versus sibling tools like migration_get_check or migration_list_checks. It lacks explicit alternatives or exclusions, relying on inference from the sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_record_approvalA
Bind a human approval to a check result: a row for exactly (stepId, checkHash). Every return of the step to the mapped state voids it, so an agent cannot commit the import after a human approved a different result.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| stepId | Yes | ||
| checkHash | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses a critical behavioral trait: the approval is voided when the step returns to the mapped state, and it warns that an agent cannot commit the import after a human approved a different result. This is valuable behavioral context beyond the schema. However, it doesn't mention idempotency behavior, error conditions, or what happens on duplicate calls, though the idempotencyKey parameter hints at that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The first sentence states the core action and key constraint; the second sentence discloses the critical invalidation behavior. Every word earns its place, and the most important behavioral warning is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with no output schema and no annotations, the description covers the essential semantics: what the tool does, the unique key, and the critical voiding behavior. It doesn't explain the full lifecycle (e.g., how to check if an approval is still valid, or what happens after binding), but the description is sufficient for an agent to understand the core operation and its main caveat. The lack of output schema means return values are unknown, but the description's focus on behavior partially compensates.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the semantic meaning of stepId and checkHash as the composite key ('a row for exactly (stepId, checkHash)'), which adds meaning. However, it doesn't explain workspaceId, planId, or idempotencyKey beyond what their names imply. The description partially compensates but leaves several parameters to be inferred from their names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Bind a human approval to a check result') and identifies the unique key ('a row for exactly (stepId, checkHash)'). It clearly distinguishes this from generic approval tools like approve_drafted_action or migration_commit_step by focusing on the binding of approval to a specific check result. However, it doesn't explicitly name a sibling tool to differentiate from, and the term 'check result' is somewhat abstract without referencing the migration check context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when a human approval needs to be bound to a specific check result, and it warns that 'every return of the step to the mapped state voids it', which is a clear condition about when the binding becomes invalid. It doesn't explicitly state when NOT to use it or name alternatives like migration_waive_control or migration_commit_step, but the context of migration checks is clear from the sibling tools and the description's language.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_rollback_stepC
Reverse a committed step through A02 reversing entries and archive master rows through their owning verbs. Never a delete: posted entries are append-only. The step returns to the mapped state and the approval is voided. After any commit this verb requires commit_migration.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| stepId | Yes | ||
| confirmed | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description must carry the behavioral burden. It discloses that the tool is not a delete (entries are append-only), archives master rows, returns to mapped state, and voids approval. However, it omits details like reversibility, permission requirements, or side effects on related data. The ambiguous requirement about commit_migration further detracts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, but the final sentence about commit_migration is unclear and the use of domain jargon like 'A02 reversing entries' without explanation reduces clarity. The structure is front-loaded but not perfectly organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a rollback operation with five parameters and no output schema, the description is insufficient. It lacks parameter explanations, prerequisites, and details about the rollback process. The agent would need additional context to call this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description provides no information about the five parameters (workspaceId, planId, stepId, idempotencyKey, confirmed) or their meanings. An agent has no guidance on what values to supply, making this a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reverses a committed step, using specific mechanisms (A02 reversing entries, archiving master rows) and explicitly disambiguates from a delete operation. It also specifies the outcome: returning to mapped state and voiding approval. This makes the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not clearly indicate when to use this tool versus alternatives like migration_abandon_plan or migration_commit_step. It mentions 'After any commit this verb requires commit_migration,' which is ambiguous and fails to give actionable guidance. There is no explicit statement of when this is the appropriate tool or what conditions must be met.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_save_map_templateA
Save a plan's finished maps as a Zuordnungsvorlage for this operator, reusable across client workspaces. Balances and every other client figure are stripped: a template carries only source labels, source account numbers and target ids.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| kinds | Yes | ||
| planId | Yes | ||
| workspaceId | Yes | ||
| sourceSystem | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses a key transformation: balances and every other client figure are stripped, so the template only contains source labels, source account numbers, and target ids. It also notes reusability across workspaces. This is valuable, though it does not cover side effects like overwriting existing templates, idempotency behavior, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose, followed by a clarifying statement about data stripping. Every word earns its place, and the structure is efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 6 required parameters, no output schema, and no annotations, the description is insufficient for correct invocation. It does not explain return values, error handling, parameter formats, or relationships (e.g., what 'kinds' represents). The core concept is clear, but an agent cannot reliably construct a valid call without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the 6 required parameters. It does not explain any parameter individually. The terms 'workspaceId', 'planId', 'name', 'sourceSystem', 'kinds', and 'idempotencyKey' are left to the agent's inference. The description gives high-level context but fails to clarify what each parameter means or how they interact, leaving a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: saving a plan's finished maps as a Zuordnungsvorlage (template) for the operator. It specifies the resource (plan's maps), the scope (operator, reusable across client workspaces), and what is excluded (balances and client figures). This distinguishes it from siblings like migration_apply_map_template and migration_list_map_templates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a specific use case: after a plan is finished, save the maps as a reusable template. However, it does not explicitly state when to use this tool versus alternatives, such as migration_set_map or migration_apply_map_template, nor does it provide conditions for when not to use it. The context is clear but lacks explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_set_manifestA
Instantiate the export-completeness manifest for a plan from its source-system guide, one item per guide row at status open. Idempotent on its key; an existing manifest is not reset (recorded statuses survive), only sourceAccessUntil is refreshed. sourceAccessUntil is the deletion-clock deadline, a contract fact to verify against your own terms.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| sourceAccessUntil | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It explicitly discloses idempotency, non-reset behavior, that existing recorded statuses survive, that only sourceAccessUntil is refreshed, and explains the meaning of sourceAccessUntil as a deletion-clock deadline. This is strong transparency for a stateful operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no wasted words. The first sentence states the core action, the second explains idempotency semantics, and the third clarifies a critical parameter meaning. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a stateful migration tool with no annotations and no output schema, the description covers the operation's purpose, its source input, its idempotency behavior, and the semantics of the key contract parameter. It could additionally mention how to retrieve the resulting manifest, but the description is otherwise complete enough for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaningful semantics for sourceAccessUntil as a deletion-clock deadline and for idempotencyKey through the idempotency statement, but it does not clarify planId or workspaceId beyond their names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the operation: instantiate an export-completeness manifest from a source-system guide, one item per open guide row. It is distinguishable from siblings like migration_set_manifest_item and migration_get_manifest, though it does not explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied through the description: it is the tool for creating the manifest from the source-system guide, and the idempotency note implies safe re-invocation. However, it does not explicitly state when to prefer this over related manifest tools or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_set_manifest_itemA
Record one manifest item: status (open, exported, not_used, blocked), the E00 fileId(s), a row count, a date range and a note. A cross-workspace fileId refuses (H-TENANT); a date range whose end precedes its start refuses. Idempotent on its key. not_used drops the item from the completeness denominator.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| dateTo | No | ||
| itemId | Yes | ||
| planId | Yes | ||
| status | Yes | ||
| fileIds | No | ||
| dateFrom | No | ||
| rowCount | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it discloses that a cross-workspace fileId is refused with H-TENANT, that inverted date ranges are refused, that the operation is idempotent on its key, and that not_used affects completeness calculations. It does not discuss authentication, permissions, or response behavior, but the stated validation and idempotency semantics provide strong transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences: the first states the resource and fields, the second states validation failures, and the third adds idempotency and denominator semantics. Every sentence earns its place, and the most important operational facts are front-loaded. There is no filler or repetition of schema structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter tool with no schema coverage and no output schema, the description covers the core workflow: what is recorded, which values are valid, two important refusal cases, idempotency, and the completeness impact of not_used. It omits return values and the semantic meaning of required identifiers, but those are contextual gaps rather than fatal ones. The richness of the behavioral detail makes this largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: it defines statuses, clarifies fileIds as E00 identifiers, identifies rowCount, date range (dateFrom/dateTo), note, and names idempotencyKey as the idempotency key. However, it does not explain workspaceId, planId, or itemId beyond their presence, nor the date string format. Still, it adds significant semantic meaning beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Record one manifest item') and enumerates the resource and fields it writes: status values, fileId(s), row count, date range, note. It clearly distinguishes its granularity from sibling migration_set_manifest by focusing on a single item. It also states key behavioral constraints, making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage through 'Record one manifest item' and the provided statuses and constraints, but it does not explicitly say when to prefer this over migration_set_manifest or related manifest siblings. There is no explicit when-to-use/when-not-to-use guidance, only implied scope. The behavioral error conditions help the caller avoid misuse but not choose among alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_set_mapB
Persist a migration map for a plan. Every target is validated against the live chart (A01) and tax codes (A05) before anything is written; one unknown target rejects the whole map, and overlapping tax validity windows are refused naming both.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | ||
| planId | Yes | ||
| entries | Yes | ||
| ruleDefault | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description carries full burden and does well by disclosing atomicity: 'one unknown target rejects the whole map' and the refusal of overlapping tax validity windows. It also clarifies that validation happens before writes. This gives the agent a clear model of the operation's failure modes and preconditions, though it could add return behavior or idempotency semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the primary purpose and followed by validation constraints. No filler words, and the structure is easy to parse. The cryptic 'A01' and 'A05' might reduce clarity, but they are domain-specific and the message remains efficiently organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a migration map with six parameters obsolete and no output schema, the description is far from complete. It does not explain the structure of a migration map, the valid kinds, how entries are formatted, what the idempotency key is for, or what response the agent can expect. The atomicity and validation rules are useful but insufficient for an agent to construct a correct call without additional documentation or examples.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description provides no parameter details. The six parameters (workspaceId, planId, kind, entries, idempotencyKey, ruleDefault) are completely unexplained. The description only vaguely references 'targets' but never maps that to the entries parameter. An agent cannot infer the required structure of entries or the meaning of kind from this definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with a specific verb and resource: 'Persist a migration map for a plan.' It clearly distinguishes this from sibling tools like migration_get_map or migration_suggest_map by describing the write operation and the validation behavior. The mention of 'Every target is validated against the live chart (A01) and tax codes (A05)' adds domain-specific scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied: the tool persists a migration map when targets need to be validated and stored. However, no explicit guidance is given on when to choose this over alternatives such as migration_apply_map_template or migration_save_map_template, nor any exclusions. The validation rule 'before anything is written' hints at its role in the migration process but does not name other tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_set_scopeB
Scope a plan per data class, one step per included class. Defaults come from discovery; a class outside first scope is returned in unavailable[] with its owning spec named, never silently absent. Excluding a class removes its step, so it can never import by accident.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| classes | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that defaults come from discovery, that out-of-scope classes are returned in unavailable[] with the owning spec named (never silently absent), and that excluding a class removes its step to prevent accidental imports. This is strong behavioral disclosure, though it could mention response structure or side effects more explicitly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The purpose is front-loaded and the behavioral details are concise. It is appropriately sized for the content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of migration scoping, no output schema, and no annotations, the description is incomplete. It mentions unavailable[] but does not describe the full response structure or what success looks like. It does not cover prerequisites (e.g., plan must exist, discovery must have run) or error conditions. It does not explain the idempotencyKey parameter. An agent would struggle to invoke this correctly without additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining parameters. It does not. It mentions 'plan' and 'class' but never maps to the actual parameters (workspaceId, planId, classes, idempotencyKey). There is no detail on how to specify classes, what idempotencyKey is for, or how workspaceId is used. This is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Scope') and resource ('a plan per data class'), and clearly describes the behavior: one step per included class. It distinguishes itself from sibling migration tools by focusing on scoping, and the description of defaults and exclusion behavior makes its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage: you use this to set the scope of a migration plan. However, it does not explicitly state when to use it versus alternatives like migration_set_manifest or migration_set_manifest_item, nor does it give exclusions or prerequisites. The guidance is present but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_suggest_mapARead-only
Propose a migration map (column, account, tax or currency) for a plan, naming which source produced it (adapter preset, locale pack, saved template or fuzzy match). Writes nothing: the caller always confirms via migration_set_map.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | Yes | ||
| planId | Yes | ||
| headers | No | ||
| dataClass | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description aligns with the readOnlyHint annotation by saying 'Writes nothing', and adds context about the proposal nature and the confirmation step. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero waste. The core purpose and the critical 'writes nothing' behavior are front-loaded, making it immediately actionable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a proposal tool with no output schema and only three required parameters, the description provides the essential workflow but leaves the semantics of headers and dataClass unexplained. Given the parameter coverage gap, it is not fully complete for an agent to call it correctly without additional inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only touches on 'kind' (via the map types) and 'plan' (implicitly planId). The parameters headers, dataClass, and workspaceId are not explained. Since there is zero coverage, the description should compensate but does not fully explain the parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb 'propose' with a clear resource 'migration map', and enumerates the map types (column, account, tax, currency) and sources. It clearly distinguishes itself from the sibling migration_set_map by noting it writes nothing, so an agent can tell it apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states the tool writes nothing and that the caller confirms via migration_set_map, which implies the proper workflow: use this to propose, then confirm with the setter. It doesn't name other alternatives like migration_discover_source, but the confirmation path is clearly indicated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_trial_load_stepA
Trial-load a step into the plan's Testmandant (G12), through the same code path the live commit uses, differing only in target workspace. Records the row-level outcomes and writes nothing to the live books. Idempotent on its key.
| Name | Required | Description | Default |
|---|---|---|---|
| planId | Yes | ||
| stepId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full safety burden. It discloses key behaviors: it writes nothing to the live books, records row-level outcomes, is idempotent on its key, and routes to a specific target workspace. This is strong for a mutation-adjacent operation, though it doesn't mention return format or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: what it does, where it writes, and its idempotency behavior. Front-loaded with the action and target. No fluff or redundant restating of the schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 4 required parameters, no output schema, and no annotations. The description covers the critical behavioral context—safety, target workspace, idempotency—that an agent needs to call it correctly. It doesn't specify the response shape, but that is less critical for a trial action. It could have mentioned the relationship to migration_commit_step more explicitly, but as a standalone description it is largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies the role of workspaceId implicitly (target workspace = Testmandant), and mentions 'idempotent on its key' which explains the purpose of idempotencyKey. However, it does not explicitly explain planId, stepId, or how the idempotencyKey should be constructed. Despite that, the freeform context is better than nothing and gives meaning to the key parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb, 'Trial-load a step,' with a clear target resource ('the plan's Testmandant (G12)') and explicitly distinguishes it from the live commit path ('same code path... differing only in target workspace'). This makes it easy to differentiate from siblings like migration_commit_step and migration_preview_step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly communicates when to use it: when you want a trial run of a migration step against the Testmandant rather than live books. It describes behavior but does not explicitly say 'use X instead of Y' or list when not to use it. Still, the contrast with the live commit path provides clear context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
migration_waive_controlA
Set one control aside with a RECORDED reason, which becomes part of the check and of the Prüfbericht. A waiver without a reason is refused (waiver_needs_reason), empty string included; a waived control reports "waived", never "passed", and readiness renders "bereit, mit N Ausnahmen" rather than plain ready. A passed control cannot be waived.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| controlId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explains the recorded reason becomes part of the check and Prüfbericht, the error condition waiver_needs_reason for missing/empty reasons, the state change to 'waived' (never 'passed'), the readiness output change, and the precondition on passed controls. This is exceptionally transparent for a mutation tool without annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two dense sentences with zero fluff. It front-loads the primary action and packs in essential behavioral details without redundancy. Every sentence earns its place, and the structure is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (4 required params, no output schema, no annotations), the description covers many critical aspects: error conditions, state transitions, readiness effects, and a precondition. However, it does not explain the return value or the purpose of the idempotencyKey, which an agent would need to know for correct invocation. These omissions leave minor gaps, but overall it is fairly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning for 'reason' (must be non-empty, recorded) and 'controlId' (the control to waive), but leaves 'workspaceId' and 'idempotencyKey' completely unexplained. Since two of four parameters lack any semantic guidance, the description only partially compensates for the schema's lack of detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Set one control aside') and the resource ('control'), and specifies that the reason is recorded. It distinguishes this tool from sibling migration tools by focusing solely on waiving a control, with no overlap in purpose. The precondition that a passed control cannot be waived further clarifies its scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context on when to use it (to waive a control with a reason) and includes a specific exclusion ('A passed control cannot be waived'), which acts as a when-not-to-use rule. However, it does not explicitly name alternative tools or contrast with them, leaving some ambiguity in selection among the many migration siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
month_end_checklistARead-only
A month-end close checklist (period YYYY-MM): dangling drafts (A02), open debtors (A16), open creditor bills (A17) and a MWST preview (A07), each with drill-down ids. Writes nothing; FX revaluation is reported as not_available until A22 builds it.
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already declares the tool is non-mutating, and the description strengthens this with 'Writes nothing.' It also adds useful behavioral nuance beyond the annotation: FX revaluation is reported as not_available until a later step builds it, and each checklist item carries drill-down ids. This is meaningful context for an agent deciding whether the expected output exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and information-dense: it states the resource, the period format, the checklist items, the drill-down capability, the read-only nature, and the one availability caveat. No sentence is wasted, and the core purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only checklist tool with two simple parameters, the description covers the key facts an agent needs: what the checklist contains, how the period is formatted, that no writes occur, and that FX revaluation may be unavailable. There is no output schema, so a slightly more explicit return structure would help, but the description is largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no descriptions for either parameter, and the description compensates for the period parameter by specifying the YYYY-MM format. However, the required workspaceId parameter is left entirely to inference. Since schema coverage is 0%, the description does not fully carry the parameter-documentation burden, though the most non-obvious parameter is addressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a month-end close checklist scoped to a YYYY-MM period and enumerates its specific contents (dangling drafts, open debtors, open creditor bills, MWST preview). It lacks an explicit verb like 'returns or lists, and does not explicitly contrast with sibling checklist_* tools, but the resource and scope are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'month-end close checklist' and the mention that FX revaluation is not_available until A22 builds it, which hints at ordering. However, there is no explicit statement of when to use this tool versus generic checklist tools (checklist_get, checklist_start) or when not to use it, so the agent must infer the intended calling context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notifications_archiveA
Archiviere eine Benachrichtigung (US-G06.2): unread oder read wird archived (reading first is not required), reversible via the Archiv filter. Self-scoped: a foreign item answers notification_not_found; archiving an archived item is a successful state assertion.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| idempotencyKey | No | ||
| notificationId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses idempotency (archiving already archived is a successful state assertion), scoping (foreign item yields notification_not_found), and reversibility via the Archiv filter. This covers key behavioral traits beyond simple mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each with essential information: what it does, a precondition note, and behavioral details. The most important information is front-loaded, and there is no filler or repetition of schema details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity and lack of output schema, the description covers the action, scope, idempotency, error handling, and reversibility. It lacks explicit parameter descriptions but the parameter semantics are partially addressed. The absence of response format is acceptable since there is no output schema and the operation is a state change.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. While it doesn't detail each parameter (workspaceId, notificationId, idempotencyKey), it explains the state semantics that guide correct parameter use: notificationId must reference a valid, scoped item, and the operation is idempotent, implying idempotencyKey use. This adds meaning beyond variable names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (archive a notification) with the specific resource and scope. It also differentiates from siblings like notifications_mark_read and notifications_list by describing the archiving behavior and the reversible filter.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains when to use it (archiving unread or read notifications) and notes that reading first is not required, which guides an agent on precondition handling. It does not explicitly name alternatives, but the scope and state assertions implicitly separate it from mark-read tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notifications_deliverA
Stelle ein Ereignis in den Posteingang zu (the OP8 action target a G01 rule names, US-G06.1): validates the event against the single automation-event registry, resolves the recipient's notification preference (exact event row, else the * wildcard, else the built-in default: inbox on), and inserts ONE unread inbox_item, or answers delivered:false reason muted when the recipient silenced the moment (never an error). An entityKind/entityId pair links back to the source record (OP3, workspace-scoped). Nothing is ever transmitted here: outbound channels are the digest's OP4 business.
| Name | Required | Description | Default |
|---|---|---|---|
| event | Yes | ||
| userId | Yes | ||
| entityId | No | ||
| entityKind | No | ||
| workspaceId | Yes | ||
| summaryParams | No | ||
| idempotencyKey | Yes | ||
| summaryI18nKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the preference resolution chain (exact, wildcard, default), the one-item insertion constraint, the muted response with delivered:false and 'never an error', and the entityKind/entityId link. It even clarifies that nothing is transmitted outbound.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense yet efficiently structured: main action first, then validation, preference resolution, outcome, and an exclusion clause. The heavy internal jargon (OP8, G01, US-G06.1, OP3, OP4) adds noise, but every sentence carries behavioral content and the layout is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with no output schema and zero parameter documentation, the description covers the delivery flow, preference resolution, and the muted edge case. Missing are error semantics for invalid events, the purpose of idempotencyKey, and the meaning of summaryI18nKey/summaryParams. These gaps are significant but the core usage is understandable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description does explain event, userId (via recipient preference), entityKind/entityId, and workspaceId. However, it leaves summaryI18nKey, summaryParams, and idempotencyKey unexplained, which is a meaningful gap for a tool with 8 undocumented parameters and no output schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Stelle ein Ereignis in den Posteingang zu' – deliver an event to the inbox) and details the core behavior: validate, resolve preference, insert one unread inbox_item, or return muted. It clearly distinguishes this from notification list/read/archive/digest siblings by describing its delivery role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains this is the automation-rule action target (OP8/G01) and explicitly excludes outbound transmission, routing that to the digest's OP4. It could be stronger with a direct statement about when an agent should call it manually versus deferring to rule execution, but the context is largely clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notifications_listARead-only
Der Posteingang (P5, US-G06.2): the caller's own inbox items newest first, each with its query-time day grouping, plus the live unreadCount (the bell figure). Self-scoped structurally: userId must be the caller's own (forbidden otherwise). status filters unread|read|archived, bucket filters to one ISO day, savedViewId applies a stored G00 view over the inbox_item kind (its filters merge underneath explicit ones).
| Name | Required | Description | Default |
|---|---|---|---|
| bucket | No | ||
| status | No | ||
| userId | Yes | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses substantive behavior: ordering is newest-first, day grouping happens at query time, unreadCount is live, userId is enforced to be the caller's own, and savedView filters merge underneath explicit filters. This is exactly the kind of behavioral context that helps an agent predict side effects and constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but not bloated; every clause adds information about scope, ordering, grouping, unread count, or filtering. The internal codes (P5, US-G06.2, G00) are somewhat cryptic but appear to be product identifiers rather than filler. A bit more structural formatting could improve readability, but the content is efficiently packed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter tool with no output schema, the description covers most operational essentials: what is returned (inbox items, day grouping, unreadCount), ordering, self-scoping, and all meaningful filters. Gaps include the role of workspaceId, default behavior when optional filters are omitted, and any pagination or result-limit caveats. Still, the description is substantially complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the parameter semantics burden. It explains status values (unread|read|archived), bucket as a single ISO day, savedViewId as a stored filter overlay, and the userId self-scoping rule. The one gap is workspaceId, which is required but never described; bucket's exact string format is also only implied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: it lists the caller's own inbox items, ordered newest first, with day grouping and the live unreadCount. It also distinguishes itself from surrounding notification tools by emphasizing self-scoping and read-only listing behavior. Even without naming a sibling, the purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: this is the read-only inbox listing tool, structurally scoped to the caller; userId must be the caller's own. It explains how filters narrow results (status, bucket, savedViewId), which tells an agent when this tool is appropriate. It stops short of explicitly naming alternatives or adding a when-not-to-use statement, but the context is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notifications_list_preferencesARead-only
Die Benachrichtigungs-Einstellungen (P5, US-G06.3): the three wildcard rows first (stored where a row exists, the built-in default where none does: inbox on, email and push off), then every stored per-event override, plus the registry's known events so a panel renders each at its inherited default rather than an empty screen. Reading another user's preferences asserts manage_members.
| Name | Required | Description | Default |
|---|---|---|---|
| userId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses meaningful behavior: the ordering (wildcard rows first, then per-event overrides), the fallback to built-in defaults (inbox on, email and push off), the inclusion of registry-known events, and the authorization requirement (manage_members) when reading another user's preferences. No annotation contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence mixing German and English, with a long parenthetical about defaults and a trailing permission note. Every clause adds information, but the structure is clunky and not front-loaded with a clear verb; it could be split into separate sentences for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description covers the essential return semantics: wildcard rows, per-event overrides, registry-known events, and the fallback defaults, along with the authorization nuance. It does not specify exact response field names or pagination, but for a straightforward read tool the described structure is largely sufficient for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry parameter meaning. It partially explains userId by noting that reading another user's preferences asserts manage_members, but it never explicitly defines workspaceId or states that userId identifies the user whose preferences are returned. The meaning of both parameters is largely left to inference from the tool name and context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explains that the tool returns notification preference rows: wildcard defaults first, then per-event overrides, plus registry-known events for panel rendering. This clearly identifies it as a list/read operation for notification settings, though it never uses an explicit verb like 'returns' and does not name sibling tools to distinguish itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context: the results are structured so a settings panel can render each event at its inherited default instead of an empty screen. It also gives an access-control guideline for reading another user's preferences (requires manage_members). However, it does not mention when to use an alternative tool such as notifications_list or notifications_set_preference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notifications_mark_all_readA
Alle als gelesen markieren (US-G06.2): clears the caller's whole unread set in one call and answers how many rows moved. Self-scoped structurally: userId must be the caller's own (forbidden otherwise).
| Name | Required | Description | Default |
|---|---|---|---|
| userId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and discloses meaningful behavioral traits: the state change ('clears... unread set'), the response shape ('answers how many rows moved'), and the authorization rule ('forbidden otherwise'). It could additionally address reversibility, but the disclosed scope and return behavior are substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences carry purpose, scope, response, and an auth constraint with minimal waste. The German prefix partially duplicates the English explanation but functions as a recognizable label. Well front-loaded with the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the essential facts: bulk effect, row-count response, and self-scoping constraint. Missing for a mutation tool with no output schema or annotations: what workspaceId scopes, idempotency semantics, and whether the clear is reversible. Adequate but with identifiable gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds critical meaning for userId ('must be the caller's own, forbidden otherwise') but says nothing about workspaceId or idempotencyKey; the latter two are left to name inference. Partial but meaningful compensation for the most security-relevant parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb+resource combination: 'clears the caller's whole unread set in one call.' The phrase 'whole unread set' distinguishes it from single-notification siblings like notifications_mark_read, and the German label 'Alle als gelesen markieren' reinforces the bulk scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context that this is the bulk operation ('whole unread set in one call'), which implicitly contrasts with the single-item alternative notifications_mark_read among the siblings. It also states the structural constraint that userId must be the caller's own. However, it does not explicitly name an alternative or state when-not-to-use.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notifications_mark_readA
Markiere eine Benachrichtigung als gelesen (US-G06.2): unread wird read und read_at gestempelt. Self-scoped: a foreign user's item answers the same notification_not_found a nonexistent one does; an ARCHIVED item is outside the mutable set and is refused the same way. Marking an already-read item read is a successful state assertion.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| idempotencyKey | No | ||
| notificationId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses that marking an already-read item is a successful state assertion, which is a key behavioral nuance. It also explains error semantics (notification_not_found for foreign/archived items) and the self-scoping constraint, adding value beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the core action and result. It packs significant detail (edge cases, error semantics) into a compact form without unnecessary verbosity. Each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and minimal annotations, the description covers essential behavioral aspects. However, it lacks details on return values or error response structure, but for a simple state-transition tool, the description is sufficiently complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It doesn't elaborate on the parameters (workspaceId, notificationId, idempotencyKey), but the purpose is clear enough that an agent can infer their roles. The description adds minimal parameter-specific meaning, leaving some ambiguity about idempotencyKey's use.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: marking a notification as read, with explicit state transition (unread to read) and timestamp. It distinguishes itself from siblings like notifications_mark_all_read and notifications_archive by focusing on a single notification read action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on when to use this tool: for a single notification, and clarifies edge cases (foreign user's item, archived items). It doesn't explicitly mention alternative tools for different scenarios, but the context is sufficient for an agent to infer appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notifications_run_digestB
Erzeuge eine Zusammenfassung (OP4, US-G06.4): gathers the recipient's inbox items in the window whose preference opted the outbound channel (email|push) into a digest cadence, renders ONE local digest artifact, and persists the digest_run honesty record. The OSS core wires no transmitter, so the answer is always transmitted:false with the reason named (cloud_tier, or empty when the window held nothing: an empty digest is never transmitted). channel inbox answers digest_channel_required. Running another user's digest asserts manage_members.
| Name | Required | Description | Default |
|---|---|---|---|
| userId | Yes | ||
| channel | Yes | ||
| periodEnd | Yes | ||
| periodStart | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden and does disclose important behavior: it always returns transmitted:false, names the reason, never transmits empty digests, and asserts manage_members for running another user's digest. The behavior is substantive and would not be inferable from the schema alone, although the "honesty record" phrasing is vague.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the core verb "Erzeuge eine Zusammenfassung," but the mixed-language phrasing and unexplained references like "OP4, US-G06.4" and "digest_channel_required" make it less readable. It contains useful behavioral details, but they are packed awkwardly rather than structured for easy agent scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given six required parameters, no output schema, and no annotations, the description lacks essential context: return value shape, parameter format requirements, idempotencyKey behavior, and exact channel constraints. It does cover transmission behavior and permission checks, but the overall picture is incomplete for an agent to invoke the tool confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has six required parameters with 0% description coverage, so the description must compensate. It only partially explains the meaning: channel is tied to email|push, periodStart/periodEnd relate to the window, and userId relates to the recipient. It leaves workspaceId and idempotencyKey undefined and does not clarify date formats, channel values, or idempotency semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: gather the recipient's inbox items in a time window, render one local digest artifact, and persist a digest_run record. It is specific enough to distinguish the tool as the digest-run operation among the notifications_* siblings, though heavy jargon such as "OP4, US-G06.4" and "honesty record" obscure the otherwise clear verb and resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly say when to use this tool versus alternatives like notifications_deliver or notifications_list. It provides some contextual constraints—no transmitter is wired, permission assertion—but no selection guidance based on user intent or digest cadence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
notifications_set_preferenceA
Setze eine Benachrichtigungs-Einstellung (US-G06.3): upserts one (recipient, event, channel) row; event omitted writes the * wildcard row every unlisted event falls back to. channel is inbox|email|push (invalid_channel otherwise); the in-app inbox is ALWAYS instant, so any other cadence on it answers inbox_is_always_instant. Setting another user's preference asserts manage_members (an admin configuring a teammate's defaults); your own needs only membership.
| Name | Required | Description | Default |
|---|---|---|---|
| event | No | ||
| digest | No | ||
| userId | Yes | ||
| channel | Yes | ||
| enabled | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it excels: it explains the upsert semantics, the wildcard fallback when event is omitted, the valid channel values and the exact invalid_channel error, the inbox_is_always_instant error for non-instant cadences on inbox, and the manage_members permission requirement. This is far beyond what the schema alone conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and packs multiple critical behaviors into a few sentences, with the core purpose front-loaded. The ticket reference 'US-G06.3' is extraneous, and the long semicolon-separated clauses make it slightly harder to parse, but every substantive sentence adds unique value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the tool, no annotations, and no output schema, the description is quite complete: it covers core upsert semantics, wildcard behavior, channel restrictions, known errors, and permission requirements. It omits some details like the exact response shape and the full semantics of digest and idempotencyKey, but an agent can invoke the tool correctly for the primary use cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds significant meaning for event (wildcard behavior), channel (allowed enum-like values and error), userId (own vs. another user's permission model), and implicitly digest as 'cadence'. However, it does not explain workspaceId, enabled, or idempotencyKey directly, leaving a few parameters to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource ('Setze eine Benachrichtigungs-Einstellung') and immediately specifies the upsert behavior and wildcard semantics. It also names allowed channel values, error cases, and permission requirements, which sharply differentiates it from sibling tools like notifications_list_preferences or notifications_deliver.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when the tool is appropriate: configuring the recipient/event/channel preference row, including distinguishing between setting your own preference versus another user's. It does not explicitly name alternatives or say when not to use this tool, but the behaviors and permission conditions make the usage scope evident.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
onboard_clientA
Onboard a new client workspace (ein neues Mandat) in one call: mints the workspace via create_workspace with its Kontenrahmen KMU chart, seeds the tax codes and sets the VAT method when vatMethod and vatAccounting are given together, and seats the operator as accepted owner (A24). Idempotent per key: replaying returns the existing workspace, never a duplicate client.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| legalForm | No | ||
| vatMethod | No | ||
| baseCurrency | No | ||
| vatAccounting | No | ||
| idempotencyKey | Yes | ||
| fiscalYearStart | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses a critical behavior: idempotency per key ('replaying returns the existing workspace, never a duplicate client'). It also explains the sequence and the conditional VAT setup. However, it does not mention permission requirements, partial failure behavior, or the return format, which are notable omissions for a mutation tool with no annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence that front-loads the purpose and packs in essential details (sub-steps, idempotency, conditional logic). It is efficient and every phrase adds value, though the long sentence with multiple parentheticals is somewhat complex to parse on the first read.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an orchestration tool with no output schema and no annotations, the description covers the main purpose, the key behavior (idempotency), and the conditional VAT logic. It does not specify the return value, error scenarios, or what happens if only one of vatMethod/vatAccounting is provided. Overall, it gives an agent enough to understand the high-level operation but lacks some specifics for full confidence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema descriptions are 0% covered, so the description must compensate. It explicitly references vatMethod, vatAccounting, and idempotencyKey, and indirectly covers name via 'new client workspace'. But legalForm, baseCurrency, and fiscalYearStart are never mentioned, leaving the agent without guidance for those fields. This is a significant gap given the total lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description has a clear verb (onboard), a specific resource (new client workspace), and enumerates the concrete sub-steps (mints via create_workspace with Kontenrahmen KMU, seeds tax codes, sets VAT method, seats operator as owner). This distinguishes it from the many individual sibling tools like create_workspace or set_vat_method by framing it as a composite 'in one call' operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'in one call' implies this is the all-in-one onboarding path, and the idempotency note provides replay semantics, but it does not explicitly state when to prefer this over calling create_workspace, set_vat_method, etc. The conditional mention ('when vatMethod and vatAccounting are given together') gives some usage context, but no explicit exclusions or alternatives are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
override_qr_matchA
The manual decision on any queue row, applied included (US-A21.4). Name the correct invoiceId to re-point (an applied row is first corrected by an A14 reversing payment, §H-AUDIT, then re-applied to the named invoice), or an action: 'unmatch' returns the row to open (reversing first when applied), 'dismiss' marks it as not a customer payment (book it via journal entry instead). Overriding an APPLIED row always requires confirmed=true: it moves money, and the auto-apply dial never covers an override. NOT automatable (D77): reversing a settlement and re-pointing money between debtors is a judgment no stored rule may make. The row keeps the reversal chain and the audit stamp: who overrode what, and when. CONSEQUENCE: Re-points or reverses a queued match, moving money between debtors through a reversing payment.
| Name | Required | Description | Default |
|---|---|---|---|
| action | No | ||
| creditId | Yes | ||
| confirmed | No | ||
| invoiceId | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full disclosure burden. It details that overriding moves money, requires confirmed=true when applied, preserves the reversal chain and audit stamp, and is not automatable. This is comprehensive transparency about side effects and constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is substantive but somewhat verbose, including internal references like US-A21.4, §H-AUDIT, and D77 that provide little value to an AI agent. The key information is front-loaded and the final CONSEQUENCE line slightly re-states earlier content, making it efficient but not maximally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description covers decision options, constraints, prerequisites, and consequences. It does not mention return values or idempotency semantics explicitly, but the core behavior is sufficiently specified for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameters. It describes invoiceId (target invoice to re-point), action values ('unmatch' and 'dismiss'), and confirmed (required when applied). It omits creditId, workspaceId, and idempotencyKey, though these are largely inferable from context and common patterns.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a manual decision on a queue row, with specific actions: re-pointing via invoiceId, unmatch, and dismiss. It distinguishes this from the auto-apply dial, but does not explicitly name sibling tools like apply_qr_match or confirm_match, so differentiation is implied rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states this is NOT automatable and that the auto-apply dial never covers an override, implying manual use for overrides. It also specifies preconditions for applied rows (requires confirmed=true, A14 reversing payment first). It does not directly name alternative tools, but provides clear contextual guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pause_recurring_scheduleA
Pause a schedule: the tick stops selecting it, immediately. Already paused settles to the same answer; an ended schedule refuses with schedule_ended.
| Name | Required | Description | Default |
|---|---|---|---|
| scheduleId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does excellently: it discloses the immediate effect, idempotency for already-paused schedules, and the specific error ('schedule_ended') for ended schedules. This is rich behavioral context beyond what the name alone would convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler; the main action is front-loaded, and edge cases follow logically. Every sentence earns its place, making it efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Behavior is well documented, but the lack of parameter semantics and no output schema leaves some invocation context unstated. The tool is simple with two obvious params, so completeness is adequate but not full – an agent might need to infer parameter meanings from context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the meaning of workspaceId or scheduleId. The parameter names are only self-descriptive in context, so the description fails to compensate for the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Pause a schedule', and explains the immediate effect ('the tick stops selecting it, immediately'). The edge cases explicitly distinguish this from resume/end siblings by describing state-dependent behavior, so an agent can tell it apart without opening other schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context that this tool stops the tick from selecting a schedule, making it the appropriate action when the schedule should be paused. However, it does not explicitly name alternatives (resume, end) or state when not to use it, so no exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
payment_batch_transmitA
Upload a generated A18 pain.001 payment batch to the bank over EBICS (BTU). P8-gated: pass confirm=true or enable the approval dial. Submits WITHOUT the authorizing signature flag, so the bank's own out-of-channel release authorizes it and TILL never holds sole payment authority; the batch shows pending_release. NEVER marks anything paid (paid comes from the camt debit via A20). Idempotent at the order level via an intent row committed before any upload: a double call delivers nothing and returns the existing order; a crash mid-upload surfaces transmit_in_doubt and blocks retransmission until the bank's protocol resolves it. A rejection lands bank_rejected with the bank's reason (recover by regenerating in A18). No channel routed to the batch's account returns needs_bank_channel and the file path stands. Posts nothing. CONSEQUENCE: Uploads the payment batch to the bank over EBICS; the bank releases it and the upload cannot be recalled.
| Name | Required | Description | Default |
|---|---|---|---|
| batchId | Yes | ||
| confirm | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It thoroughly reveals idempotency ('a double call delivers nothing and returns the existing order'), crash behavior ('surfaces transmit_in_doubt and blocks retransmission'), rejection outcomes ('lands bank_rejected'), and the consequence that the upload cannot be recalled. This exceeds typical transparency and gives the agent complete expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary action and then layers in crucial behavioral nuances. It is lengthy but every sentence adds value, covering idempotency, error states, rejection handling, and consequences. Some redundancy exists (e.g., the 'CONSEQUENCE' section restates the upload action), but the overall density is justifiable for a high-risk financial transmission tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the operation and the absence of annotations and output schema, the description covers all critical aspects: authorization gate, idempotency, failure modes, recovery, and side effects. It even describes return states ('returns the existing order', 'shows pending_release', 'lands bank_rejected', 'returns needs_bank_channel'), making it self-sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. It explicitly explains the `confirm` parameter ('P8-gated: pass confirm=true') and indirectly covers `idempotencyKey` via the idempotency statement. However, `batchId` and `workspaceId` are not explicitly described, though their purpose is inferable from context. Partial compensation warrants a baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb, resource, and destination: 'Upload a generated A18 pain.001 payment batch to the bank over EBICS (BTU).' It clearly differentiates from related tools by stating it does not mark payments paid ('NEVER marks anything paid') and does not post anything, which distinguishes it from payment posting or marking tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes the condition for use: 'P8-gated: pass confirm=true or enable the approval dial.' It also clarifies what the tool does NOT do ('NEVER marks anything paid', 'Posts nothing'), which serves as a when-not to use it for payment confirmation tasks. The idempotency and error-handling details further guide an agent on safe invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
payroll_handoff_exportA
Export the employee master (E02) plus the mutations since the last export, as a LOCAL artifact for an external payroll provider. format is csv or json (default json). The artifact is stored via E00 and linked to a payroll_handoff record; it NEVER transmits (OP4). AHV numbers are included only when the actor holds hr.sensitive; when they are not, the artifact is produced WITHOUT them and both the artifact header and the result say so ({ahvIncluded:false, ahvExcludedReason:'missing_hr_sensitive'}), never a silent omission (revDSG Art. 6). The first export is the full master with an empty mutations section; a re-run with no changes yields mutationCount:0. Refuses with no_employees when the roster is empty. Posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations and no output schema, the description carries the full behavioral burden and does so thoroughly: it discloses storage via E00, non-transmission, AHV inclusion conditions with explicit result fields, first-run behavior, empty-rerun behavior, and the no_employees refusal. This goes well beyond a basic description and leaves little hidden behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but mostly well-organized, with the core purpose first and important edge cases and safety behavior after. Every sentence covers a distinct aspect, though the single-paragraph format and slight redundancy between 'NEVER transmits' and 'Posts nothing' keep it from being fully polished.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a payroll-export operation with sensitive data, no annotations, and no output schema, the description covers output format, storage, non-transmission, permissions-driven AHV behavior, idempotent reruns, empty-roster errors, and side effects. An agent has enough context to invoke this correctly and understand expected results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds real meaning for the format parameter ('csv or json (default json)') and partially implies idempotent behavior ('a re-run with no changes yields mutationCount:0'). However, schema coverage is 0% and the description does not meaningfully explain workspaceId or idempotencyKey, leaving part of the parameter semantics to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Export'), a precise resource ('employee master (E02) plus the mutations since the last export'), and a clear destination ('LOCAL artifact for an external payroll provider'). It also distinguishes itself from transmission-style tools by explicitly noting 'it NEVER transmits (OP4)' and 'Posts nothing.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly frames when this tool is appropriate: exporting payroll-related data to an external provider, including 'first export' vs 're-run' semantics and a failure case ('Refuses with no_employees'). It does not explicitly name sibling alternatives or exclusions, but the use case is unambiguous enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pipeline_stages_upsertA
Lege eine Phase an oder bearbeite sie (§6b): name, sort, the default probability (0 to 100, validated here, the one place a default enters), and the outcome flag (won or lost) that makes entering the stage terminal through deals_mark. Pass stageId to patch, omit it to append. The stage names are workspace data; the stage-to-status derivation stays fixed mechanism (§6b Fixed).
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| sort | No | ||
| outcome | No | ||
| stageId | No | ||
| pipelineId | Yes | ||
| probability | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses that probability is validated (0-100), outcome flags terminal status via deals_mark, and that the stage-to-status derivation is fixed. It does not mention permissions or side effects, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise (three sentences) and front-loaded with the action. It includes a §6b reference that may be opaque but is domain-specific. No redundant filler; efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an upsert with 8 params and no output schema, the description covers the main mechanics but lacks return-value expectations, error handling, and explicit mention of workspaceId/pipelineId. It is adequate for a knowledgeable user but incomplete for an autonomous agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters. It covers name, sort, probability, outcome, and stageId, but omits workspaceId, pipelineId, and idempotencyKey. While the first two are required and inferable, idempotencyKey is entirely unexplained, and sort is only listed without meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool creates or edits a pipeline stage, listing the key fields (name, sort, probability, outcome) and explaining the patch/append behavior via stageId. It differentiates from related tools like pipelines_upsert (for pipelines) and deals_mark (which consumes the outcome flag).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit usage guidance: pass stageId to patch, omit to append, and notes that probability is validated here as the only place defaults enter. It implies this is the tool for stage management but does not explicitly name alternatives or exclusions; still, the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pipelines_upsertB
Lege eine Pipeline an oder benenne sie um (§6b, the OP10 flexible surface for pipeline shape): pass pipelineId to rename, omit it to create. The stage set itself is edited per stage through pipeline_stages_upsert.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| pipelineId | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description bears the full burden of behavioral disclosure. It reveals the conditional create/rename behavior but omits other important traits: mutation side effects, permission requirements, reversibility, or the meaning of the idempotencyKey parameter. The cryptic 'flexible surface' phrase is not self-explanatory. Overall, only minimal behavioral context is given.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is relatively short, but the first sentence (German) repeats the create/rename concept already covered in the second sentence. The cryptic '§6b' and 'OP10 flexible surface' are unhelpful. It could be more concise by removing the German redundancy and jargon, though the core instruction is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 parameters, no output schema, and no annotations, the description is incomplete. It covers the create/rename logic and points to the stage tool, but fails to explain the required workspaceId, the name parameter, or idempotencyKey behavior, and does not describe any return value. An agent would struggle to construct a correct call without additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain all parameters. It only explains pipelineId (pass to rename, omit to create). The required workspaceId, name, and idempotencyKey are not described. This leaves the agent guessing about their roles and formatting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's dual purpose: create a new pipeline or rename an existing one, determined by the presence of pipelineId. It also distinguishes itself from pipeline_stages_upsert for stage editing. However, the cryptic references to '§6b' and 'OP10 flexible surface' add noise without explanation, slightly reducing clarity for an agent unfamiliar with internal jargon.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly provides usage guidance: pass pipelineId to rename, omit it to create. It also directs the agent to a sibling tool for stage editing, effectively stating when not to use this tool. This is a clear decision rule for selecting this tool over alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_amendment_applyA
Wende eine Änderung an und erzeuge die nächste PO-Version (OP14): supersedes the current active version, updates the live purchase_order + po_line to the amended values (received_qty and billed_qty on surviving lines are NEVER touched), mints version N+1 (active) with the frozen snapshot, and RE-RENDERS the outbound PO artifact carrying the new revision. P8: it returns { transmitted:false } and never emails the supplier on its own; an amended commitment reaches the supplier only through the approval dial. Applies from draft or pending_approval. Refuses qty_below_received / line_has_receipts / no_effective_change (a refusal writes zero rows). Idempotent on ROWS: a replay under the same idempotencyKey creates NO second version.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| amendmentId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and delivers extensively: it details that it supersedes the active version, updates live purchase_order and po_line, never touches received_qty/billed_qty, mints version N+1, re-renders the artifact, returns { transmitted:false }, never emails the supplier, is idempotent on rows, and writes zero rows on refusal. This is exceptionally transparent about side effects and edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and front-loaded with the primary purpose. It uses semicolons and parentheses to pack many details into one paragraph, which is efficient but could benefit from bullet points or clearer separation of concerns. Every sentence contributes value, but the density may reduce readability for an agent scanning quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers many behavioral aspects, but it is ambiguous about the return value for success and failure beyond the 'P8' mention of { transmitted:false }. It does not explicitly state what the tool returns on a successful application or on a refusal (only that refusal writes zero rows). Given there is no output schema, this is a notable gap. It also introduces 'P8' without explaining it, and does not clarify the required parameters' roles fully.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the idempotencyKey behavior ('a replay under the same idempotencyKey creates NO second version') but does not explain workspaceId, amendmentId, or actor. The purpose statement implies amendmentId refers to the amendment being applied, but no explicit parameter semantics are given. It adds some value but leaves several parameters unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Wende eine Änderung an und erzeuge die nächste PO-Version' (Apply an amendment and generate the next PO version). It specifies the action (apply amendment) and the resource (PO version), and it differentiates from siblings by being the apply step in the amendment workflow, as opposed to start, update_lines, preview, or submit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit conditions for use: 'Applies from draft or pending_approval' and lists refusal cases (qty_below_received, line_has_receipts, no_effective_change). It also explains that it never emails the supplier, which helps decide when to use it. However, it does not explicitly name alternative tools or state 'use this instead of X', so it lacks direct exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_amendment_cancelA
Verwirf eine Entwurfs- oder eingereichte Änderung (OP14): status -> cancelled, the live PO untouched, and a new amendment may then be started. An already-applied amendment cannot be cancelled (invalid_transition); a further corrective amendment is required. Idempotent.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| reason | No | ||
| amendmentId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It covers the state transition to cancelled, guarantees the live PO remains untouched, notes idempotency, and warns about the invalid_transition error for applied amendments. This is strong coverage of side effects and failure modes, though it omits details like permission requirements or return payload.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence that front-loads the primary action and outcome. It includes essential caveats (invalid_transition, idempotent) without verbosity. While it mixes German and English, it remains readable and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a workflow-specific cancel operation with no output schema and no annotations, the description covers the core state transition and failure case but omits parameter explanations and return details. It also does not address what happens on success (e.g., return object) or any additional side effects beyond the state change. Given the moderate complexity, this is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 5 parameters with 0% description coverage, and the description does not explain any of them beyond the implicit use of amendmentId. It mentions idempotency, which hints at idempotencyKey, but actor and reason are left entirely unaddressed. Given the lack of schema descriptions, this is a significant gap that the description fails to compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (discard an amendment), the exact resource (draft or submitted amendment, OP14), and the outcome (status -> cancelled, live PO untouched). It distinguishes from related tools by explicitly limiting to draft/submitted states and noting that applied amendments cannot be cancelled, which differentiates it from po_amendment_reject or others.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly indicates when to use it: to discard a draft or submitted amendment. It also specifies a condition for non-use (already-applied amendments cannot be cancelled, requiring a corrective amendment instead). It does not explicitly name sibling alternatives, but the state-based constraints provide clear routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_amendment_previewARead-only
Zeige die Auswirkung einer Änderung, bevor sie angewandt wird (OP14, P5, pure): per-line before/after qty and unit price, the committed-value delta in Rappen, and the violation list (qty_below_received when a line would fall below its received quantity, line_has_receipts when a received line would be removed). Nothing is written. The same impact the apply enforces, so preview and apply never disagree.
| Name | Required | Description | Default |
|---|---|---|---|
| amendmentId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description states 'Nothing is written' and 'pure,' directly confirming no side effects. It also discloses exactly what the preview computes, including violation semantics and the consistency guarantee with apply. This is strong behavioral disclosure for a read-only tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then efficiently packs output details, side-effect guarantees, and consistency semantics. The inline tags 'OP14, P5, pure' are opaque and add little for an agent, but the overall structure remains focused and information-dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With only two required parameters and no output schema, the description supplies the key return values and behavior needed for an agent to understand the tool. It omits details such as exact response shape or whether the amendment must be in a draft state, but these are minor given the tool's read-only preview role.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain workspaceId or amendmentId or how they relate to the amendment being previewed. The parameter names are fairly self-explanatory, which prevents a score of 1, but no parameter-level guidance is added.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it shows the impact of an amendment before it is applied, with detailed outputs like per-line before/after quantities, unit price, committed-value delta, and violations. It clearly distinguishes itself from the application step and sibling tools such as po_amendment_apply.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the intended use explicit: call this to see the effect before applying and trust that it matches the eventual apply behavior. It contrasts with the apply tool by stating 'Nothing is written' and 'preview and apply never disagree,' though it does not enumerate exclusions or compare with all other amendment siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_amendment_rejectA
Lehne eine Änderung als Genehmiger ab (OP14): status -> rejected with a mandatory reason, the live PO untouched. Allowed from draft or pending_approval; an applied amendment cannot be rejected (invalid_transition). A missing reason is invalid_input. Idempotent.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| reason | Yes | ||
| amendmentId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full responsibility. It discloses the status change, that the live PO is untouched (side effect), mandatory reason, valid state preconditions, two specific error codes (invalid_transition, invalid_input), and idempotency. This is comprehensive and leaves no ambiguity about behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence plus the word 'Idempotent.' It front-loads the action and all key constraints without waste. Every clause carries essential information, making it highly scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the state transition, allowed states, error conditions, side effects, and idempotency—sufficient for an agent to call the tool correctly. It does not describe the response format, but no output schema is present, and for a reject operation that is not critical. Minor omission: no explicit mention of what a successful call returns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 0% description coverage, so the description must compensate. It states 'reason' is mandatory and mentions idempotency (hinting at idempotencyKey). However, it does not explain workspaceId, amendmentId, or actor beyond their obvious roles. The added value over schema names is modest but non-zero.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: reject an amendment as approver, with the effect 'status -> rejected' and that the live PO remains untouched. It distinguishes itself from sibling operations like po_amendment_apply or po_amendment_submit by specifying this is the reject path (OP14) and by naming the exact state transition.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly lists the allowed source states (draft or pending_approval) and the invalid state (applied amendment) with the resulting error (invalid_transition). While it doesn't name alternative tools, the state-based guidance is concrete and tells the agent when this tool is appropriate. A small gap is not pointing to a sibling for approving or canceling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_amendment_startA
Beginne eine Änderung an einer gesendeten oder teilweise erhaltenen Bestellung (OP14): opens a DRAFT amendment linked to the current active version. No live PO data is mutated yet. Materialises version 1 first if needed (US-I01.1). A PO that is draft/closed/cancelled is refused (invalid_transition); a PO with no open quantity left is nothing_open; a second amendment while one is already draft/pending_approval is amendment_in_progress (one open amendment per PO). Off the money path (P3); idempotent.
| Name | Required | Description | Default |
|---|---|---|---|
| poId | Yes | ||
| actor | No | ||
| reason | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses key behaviors: no live mutation, version materialisation, idempotency, and the exact refusal conditions. This is a high level of transparency for a mutation-like tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but front-loaded with the core action and safety note. It packs error conditions efficiently, though it reads as a long single sentence. It is not wasteful, but structuring into separate points would improve readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
It covers the start action and error states well, but omits the return value (likely a draft ID) and does not explain the purpose of actor/reason. Given no output schema, the description should indicate what the caller receives on success. It is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it provides no parameter-level guidance. It only implicitly clarifies poId (the PO to amend) and omits meaning for actor, reason, and idempotencyKey. The description adds minimal value for parameter usage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (opens) and resource (a DRAFT amendment) tied to a sent/partially received PO. It distinguishes from sibling amendment tools by being the start step and explicitly notes no live data mutation, so an agent can tell it apart from po_amendment_submit/apply/cancel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit conditions for when the tool is valid: PO must be sent/partially received, not draft/closed/cancelled, have open quantity, and no existing open amendment. It implies it is the entry point to the amendment flow, though it does not name alternative tools directly. The error conditions (invalid_transition, nothing_open, amendment_in_progress) give clear guidance on when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_amendment_submitA
Reiche eine Änderung zur Genehmigung ein (OP14, draft -> pending_approval): for a workspace that gates material amendments before they become the live commitment. Requires a valid, effective impact (a zero-delta amendment is no_effective_change; a blocking violation is returned by its code). Only a DRAFT amendment can be submitted (invalid_transition). Idempotent.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| amendmentId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the state transition, idempotency, and specific error conditions (no_effective_change, invalid_transition). This is useful behavioral detail, though it does not describe the success response or downstream approval effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core action and transition. Every clause conveys a constraint or behavior. It is slightly dense due to the German/English mix and multiple parenthetical error codes, but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main preconditions, state transition, error cases, and idempotency. However, there is no output schema and no mention of what the successful submission returns or what happens after approval is granted. This leaves an agent without a clear model of the operation's full outcome.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the four parameters. It only indirectly suggests meaning for workspaceId and amendmentId via the workspace/amendment context, and never defines actor or idempotencyKey. Parameter names are self-explanatory, but the description itself adds little beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('submit for approval'), a specific resource (a material amendment), and a state transition (OP14, draft -> pending_approval). This clearly separates it from sibling tools like po_amendment_apply, po_amendment_cancel, and po_amendment_reject.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context: use this only in workspaces that gate material amendments before they become live, and only when the amendment is in DRAFT. It also states that zero-delta amendments or invalid transitions are rejected. It does not explicitly name alternatives, but the preconditions are clear enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_amendment_update_linesA
Setze die Änderungspositionen einer Entwurfs-Änderung (OP14): replaces the amendments set of change operations. Each op is change (edit a live lines qty/unitPriceRappen/description; a field left out keeps the live value), add (a new line: itemId or free-text description, qty > 0, unitPriceRappen >= 0), or remove (drop an open line). Only a DRAFT amendment is editable (invalid_transition otherwise). A poLineId that is not a live line of this PO is not_found; an unknown itemId is invalid_reference; qty <= 0 is invalid_qty. Validated as a batch: a refusal writes zero rows.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| changes | No | ||
| amendmentId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that the operation replaces the existing set of change operations, provides validation error semantics (invalid_transition, not_found, invalid_reference, invalid_qty), and states that validation is batch-atomic ('a refusal writes zero rows'). This is substantial behavioral disclosure, though it omits details like permission requirements or the response structure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph in mixed German/English. It front-loads the core purpose, then details operation types and validation rules, with no filler. It could benefit from bullet points for readability, but it is still concise and every sentence contributes value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the description covers many important aspects: the set-replacement behavior, error cases, and atomicity. It does not explain the response format (no output schema exists, so this is acceptable per rubric) and omits details like how to specify a 'remove' target (presumably by poLineId). Still, it provides enough for an agent to make a well-formed call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It thoroughly explains the 'changes' parameter, describing each operation type and its constraints (e.g., qty > 0, unitPriceRappen >= 0, itemId or free-text description). However, it does not explain the 'actor', 'idempotencyKey', or 'workspaceId' parameters, though these may be conventional. The complex parameter is well covered, but not all parameters are.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear statement that the tool replaces the amendment's set of change operations, which is specific about the action and resource. The breakdown of each operation type (change, add, remove) further clarifies its purpose. The name itself also helps, but the description goes beyond the name to define the exact role among the po_amendment_* siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states that only a DRAFT amendment is editable, which is a key condition for when to use the tool. It does not explicitly name alternative tools, but the context of 'updating lines' is clear enough in relation to siblings like po_amendment_start, po_amendment_preview, etc. This is better than a vague hint but not a full routing guide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_cancelA
Storniere eine Bestellung (draft|sent -> cancelled): ONLY while nothing has been received. A cancelled PO keeps its rows (H-AUDIT, documents are trail, not trash). Cancelling a PO with any receipt is refused with has_receipts (use po_close_short instead); a non-draft/sent PO is refused (invalid_transition).
| Name | Required | Description | Default |
|---|---|---|---|
| poId | Yes | ||
| actor | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Even though no annotations are provided, the description discloses crucial behavioral traits: the allowed state transition, the precondition of no receipts, and that cancelled POs retain their rows as an audit trail ('documents are trail, not trash'). It also specifies error codes (has_receipts, invalid_transition) and the outcome of refusing the operation. This goes well beyond a simple action statement and makes the tool's behavior transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense sentence with parenthetical clarifications. Every element earns its place: the state transition, the precondition, the row-retention effect, the refusal codes, and the alternative tool. There is no redundant wording, and the most critical constraint ('nothing has been received') is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a cancellation operation with no output schema, the description covers the essential decision-making context: preconditions, effects on data, error conditions, and alternatives. However, it does not explain the parameters (especially idempotencyKey) or any side effects like notifications or audit log entries beyond the mention of H-AUDIT. The action is relatively simple, so the description is largely complete, but the parameter gap keeps it from being fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has four parameters (poId, actor, workspaceId, idempotencyKey) with zero descriptions, and the description provides no additional meaning about any of them. It does not explain that poId identifies the purchase order to cancel, what actor or workspaceId refer to, or the role of idempotencyKey. With schema description coverage at 0%, the description fails to compensate for the complete lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Storniere' (cancel) and the resource 'Bestellung' (purchase order), and specifies the exact state transition (draft|sent -> cancelled). It also distinguishes this tool from po_close_short by noting that cancellation is refused when receipts existhol that po_close_short is for that case. The purpose is unambiguous and well-separated from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool ('ONLY while nothing has been received') and when not to use it (if there are receipts, use po_close_short; if the PO is not in draft or sent state, it is refused with invalid_transition). This gives clear when/when-not guidance and names the alternative tool, leaving no ambiguity for an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_close_shortA
Schliesse die Restmenge einer Bestellung (sent -> closed): waives the remaining open quantity of a partially received PO. The backorder is explicitly waived, never silently dropped; the PO keeps all its rows (H-AUDIT). A non-sent PO is refused (invalid_transition).
| Name | Required | Description | Default |
|---|---|---|---|
| poId | Yes | ||
| actor | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses that the backorder is explicitly waived, never silently dropped, that the PO keeps all rows (H-AUDIT), and that non-sent POs are refused. It doesn't mention permissions or reversibility, but the core side effects and failure mode are clearly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the first clause gives the action and transition, the second clarifies the waiver semantics, and the third states the refusal condition. No wasted words, every sentence adds essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a state-transition tool without output schema or annotations, the description covers the key prerequisites (partially received, sent state), the effect on rows, and the failure mode. It doesn't explain response format, but that is not required without an output schema. Minor gaps remain around parameter usage, which is covered in that dimension.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it never explains the parameters (poId, workspaceId, actor, idempotencyKey). An agent must infer that poId identifies the PO and workspaceId scopes the operation; this is a significant gap for a tool with 4 parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'waives the remaining open quantity of a partially received PO' with an explicit state transition (sent -> closed). This clearly differentiates it from sibling tools like po_cancel or po_revise by focusing on partially received POs and backorder waiver.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: it is for partially received POs where the remaining quantity is to be waived, and it explicitly says a non-sent PO is refused (invalid_transition). However, it does not directly compare to alternatives like po_cancel or po_revise, leaving the choice partly to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_getARead-only
Lies eine Bestellung mit Positionen (bestellt / erhalten / verrechnet / offene Menge), ihren Wareneingängen und den 3-Way-Match-Sätzen. The one read that shows the whole PO -> receipt -> match chain for a single order.
| Name | Required | Description | Default |
|---|---|---|---|
| poId | Yes | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes that this is a safe read operation. The description adds meaningful behavioral detail: it reveals exactly what data the tool returns, including positional quantities, goods receipts, and 3-way-match sets, and scopes it to a single order. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded, with the first sentence listing the concrete data blocks and the second sentence summarizing the end-to-end chain. There is some redundancy between the German enumeration and the English chain summary, but both sentences contribute useful framing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-order read, the description adequately names the main contents and the chain scope. However, it leaves savedViewId unexplained and, with no output schema, offers no indication of result shape or pagination behavior. It is workable but not fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description does not explain workspaceId, poId, or savedViewId. The context only implies poId identifies the order, while savedViewId remains entirely unexplained. The description fails to compensate for the undocumented parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a read operation ('Lies') on a single purchase order and enumerates the included data: positions with ordered/received/billed/open quantities, goods receipts, and 3-way-match sets. The phrase 'the whole PO -> receipt -> match chain' differentiates it from more narrowly scoped PO or match tools, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this is the read to use when the complete PO-to-receipt-to-match chain for a single order is needed. It does not explicitly list alternatives or conditions to avoid, but 'the one read that shows the whole ... chain' provides a strong selection signal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_listARead-only
Liste die Bestellungen (P5): filter by status or supplier. Each row carries the number, status, currency, totals and expected date. savedViewId support rides G00`s saved-view seam on the po customization surface.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| supplierContactId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description does not contradict that. It adds that each row carries specific fields and that savedViewId is supported, providing some behavioral context beyond the annotation. However, it does not disclose pagination or sorting behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, concise and front-loaded with the primary action. The second sentence introduces savedViewId with somewhat cryptic language ('G00`s saved-view seam'), which may reduce clarity, but overall it is brief.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides the key filter options and lists the fields returned, but lacks details on pagination, sorting, or the required workspaceId. Given that there is no output schema and only readOnlyHint annotation, the description could be more complete for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameter roles. It explains status and supplierContactId as filters, and mentions savedViewId support. It does not explain workspaceId, which is required, but that is often self-evident. This adds value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it lists purchase orders (P5) and specifies available filters (status, supplier). It distinguishes from other PO operations by indicating a list function, though it does not explicitly name sibling alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies usage for listing orders with optional filters, but does not explicitly state when to use this vs. other PO tools (e.g., po_get for a single order). No exclusion criteria are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_open_linesARead-only
Liste die offenen Mengen / Lieferrückstände (P5): every PO line with received_qty < qty on a SENT PO, with open_qty = qty - received_qty. An agent polls this to chase a supplier after a partial delivery. An empty list is not an error.
| Name | Required | Description | Default |
|---|---|---|---|
| savedViewId | No | ||
| workspaceId | Yes | ||
| supplierContactId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations include readOnlyHint: true, which is consistent. The description adds valuable behavior context: 'An empty list is not an error' and explains how open_qty is computed. This goes beyond the annotation by clarifying edge-case semantics and the data source condition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero fluff. The core definition is front-loaded, followed by a practical use case and an edge-case note. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core behavior (list open lines, computation, empty list handling) but omits explanation of the optional parameters, which are not self-explanatory from the schema. For a read-only list tool, this is a moderate gap; an agent might misuse savedViewId or supplierContactId without guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%: the description does not mention any parameters (savedViewId, supplierContactId) beyond workspaceId being implicitly required. Since the schema has no descriptions, the tool description must compensate, but it doesn't explain what these optional parameters do or how they filter results. This is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: lists open quantities / delivery backlogs on PO lines. It precisely defines the filter (received_qty < qty on SENT PO) and the computed open_qty. This clearly distinguishes it from siblings like po_list, po_get, and po_revise, which have different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states a use case: 'An agent polls this to chase a supplier after a partial delivery.' This gives clear context for when to invoke it. It does not mention alternatives or when not to use it, but the specific scenario is sufficient guidance for an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_reviseA
Revidiere eine gesendete Bestellung (sent -> draft, die Revise-Kante): snapshots the current header + lines into po_revision (append-only, H-AUDIT), increments revision, and re-opens the PO as a draft for editing via po_upsert. received_qty and billed_qty survive untouched (they belong to receipts/matches). The re-send rides the P8-gated po_send, so a revision can never leak to the supplier without the dial. Only a SENT PO is revisable (draft/received/closed/cancelled -> invalid_transition). Idempotent: a replay yields one snapshot + one increment.
| Name | Required | Description | Default |
|---|---|---|---|
| poId | Yes | ||
| actor | No | ||
| reason | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so thoroughly. It discloses the snapshot/append-only audit behavior (po_revision, H-AUDIT), the revision increment, the preservation of received_qty/billed_qty, the P8-gated re-send safety property, and idempotency semantics. This is rich behavioral context beyond what any schema could convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: state transition, audit behavior, quantity preservation, re-send gating, valid-state constraint, and idempotency. It is front-loaded with the core action and uses compact technical notation (sent -> draft, H-AUDIT, P8-gated) that is efficient for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers the essential behavioral contract: what changes, what is preserved, what is recorded, when it is valid, and what happens on replay. The only minor gap is explicit parameter semantics, but the operation's context makes the required parameters clear. The description is complete enough for an agent to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the domain meaning of the operation but does not explicitly map the five parameters (workspaceId, poId, actor, reason, idempotencyKey) to their roles. However, the idempotency statement and the state-transition context make the purpose of poId, workspaceId, and idempotencyKey inferable. The description adds meaningful context but leaves some parameter-level detail to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Revidiere eine gesendete Bestellung') and immediately states the state transition (sent -> draft). It names the sibling po_upsert as the follow-up editing tool and po_send as the re-send path, which distinguishes it from the many other PO tools in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when the tool is valid ('Only a SENT PO is revisable') and lists the invalid states (draft/received/closed/cancelled -> invalid_transition). It also names the alternative editing path (po_upsert) and the re-send path (po_send), giving clear routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
portal_grant_createA
Gib einem Kunden Portal-Zugang frei: mint a scoped, expiring, tokened grant for a C00 contact. scopes is an explicit entity list ([{kind:"invoice"|"quote"|"document", id}] or {kind:"all_invoices"}); every id must be THIS contact's own record (invalid_scope otherwise, H-TENANT + per-contact fence). The token is CSPRNG >=256-bit; only its SHA-256 hash is stored, and the row carries NO usable link. The one-time link is returned once as tokenOnce/localLink and is never re-derivable. expiresAt in the past is refused; longer than the 90-day max is CLAMPED and flagged (clamped:true). Created draft (P8): handing the link over is portal_grant_send. Posts nothing (OP4: mints the local artifact and stops, hosted:false/reason:cloud_tier).
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| actor | No | ||
| scopes | Yes | ||
| contactId | Yes | ||
| expiresAt | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so thoroughly. It reveals token generation (CSPRNG >=256-bit), hash-only storage, no usable link in the row, one-time link non-re-derivability, expiry clamping with a flag, draft creation, and the fact that no post occurs (OP4, hosted:false). It also covers scope validation rules. This is exceptionally transparent and leaves no major behavioral surprises.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but appropriately so given the complexity. It is front-loaded with the core purpose and then adds essential behavioral details. Each clause contributes valuable information without redundancy. It is longer than a simple one-liner, but every sentence earns its place, covering security, scope validation, expiry behavior, and workflow context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description must convey return values, and it does: tokenOnce/localLink is returned once, and clamped:true is flagged. It also explains the lifecycle (draft, then send) and security implications. It does not enumerate the full response structure or list error codes beyond invalid_scope, but it covers the essentials an agent needs to call the tool correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate, and it does for the critical parameters. It explains scopes in detail (entity list, kinds, all_invoices), expiresAt (past refused, max 90 days clamped), and contactId (must be this contact's own record). It does not explain workspaceId, kind, actor, or idempotencyKey, but those are less central or optional. It adds significant meaning beyond the schema for the required inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('mint a scoped, expiring, tokened grant') and names the exact resource (a C00 contact). It clearly distinguishes itself from the sibling portal_grant_send by noting that creation happens first and sending is a separate step. The verb 'Gib einem Kunden Portal-Zugang frei' plus the detailed mechanics leave no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly points to portal_grant_send as the next step ('handing the link over is portal_grant_send'), which tells the agent when to use this tool versus that one. It does not explicitly enumerate when not to use it or contrast with revoke/list, but the creation-then-send workflow is clear. This is strong usage guidance, though not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
portal_grant_listARead-only
Liste die Portal-Freigaben (P5): every grant, optionally per contact, each with its derived status (draft/active/revoked/expired) and its hosted:false truth (the same shape an agent and the panel both see). savedViewId applies a saved view (G00): its stored filters merge underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| contactId | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description adds that it returns derived statuses (draft/active/revoked/expired) and the hosted:false truth, and that the shape matches what agents and the panel see. This provides useful behavioral context about the output without contradicting annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no redundancy. It front-loads the main purpose and then explains the optional filter behavior. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the output shape (status, hosted:false truth), the same shape as agent/panel, and saved view behavior. It does not mention pagination or sorting, but for a listing tool with readOnlyHint, it is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explains the semantics for contactId (optionally per contact) and savedViewId (applies saved view with filter merging), but does not mention workspaceId, which is the only required parameter. Since schema coverage is 0%, the description compensates partially but not fully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it lists portal grants (P5), optionally per contact, with derived status and hosted:false truth. The verb 'Liste' and resource 'Portal-Freigaben' are specific, and it distinguishes itself from sibling tools like vendor_portal_grants_list by focusing on the portal grants context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies it is a general listing tool for portal grants, with optional contact filtering and saved view application. However, it does not explicitly state when to use this tool versus alternatives like vendor_portal_grants_list, nor does it provide exclusions or specific conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
portal_grant_revokeA
Widerrufe einen Portal-Zugang: stamps revoked_at so any resolve of the token now denies. The row is NEVER deleted (the grant history is the revDSG access trail). Revoking an already-revoked grant is a no-op returning the original state. Not destructive: it revokes, it does not delete.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| grantId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations supplied, the description carries the full behavioral burden and does so thoroughly. It discloses that the operation is non-destructive, permanently marks revoked_at, and returns the original state for already-revoked grants, which is exactly the kind of side-effect information an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded, with the primary action in the first word and key consequences following. It is slightly redundant ('row is NEVER deleted' and 'revokes, does not delete'), but every sentence adds relevant behavioral information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter tool with no annotations and no output schema, the description covers the core semantics unusually well: effect, non-deletion, idempotency, and denial on later resolves. It falls short only on documenting parameter roles and the expected success return value.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, and the description does not explain any individual parameter (grantId, workspaceId, actor, idempotencyKey). While the names are somewhat self-explanatory, the description does not compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action—revoke a portal access grant—and explains the core effect (stamps revoked_at, so token resolution now denies). It is clear about what the tool does, but it does not explicitly differentiate itself from sibling revoke tools such as vendor_portal_revoke.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when it should be used (when a portal access grant needs revoking) and adds useful contextual notes (revoking an already-revoked grant is a no-op, and the operation is non-destructive). However, it provides no explicit when-not-to-use guidance or named alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
portal_grant_sendA
Sende den Portal-Link (P8 outbound, draft-gated): hands the grant link over and activates it (draft -> active). Confirm-gated (needs_confirmation without confirmed:true, for humans and agents alike). OP4 boundary: no transport in the MIT core, so it degrades honestly to sent:false/reason:cloud_tier and the token is NOT re-derivable here (the operator received the one-time link at create). A revoked or expired grant is refused before any write.
| Name | Required | Description | Default |
|---|---|---|---|
| grantId | Yes | ||
| confirmed | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It discloses the confirmation requirement, the honest degradation to sent:false/reason:cloud_tier, the non-re-derivability of the token, and the refusal before any write for revoked/expired grants. These are critical behavioral traits that affect outcomes and error handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient, with each sentence adding unique value. The main purpose is front-loaded, followed by confirmation, degradation, and refusal conditions. It avoids repetition and is appropriately sized for the complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema and annotations, the description covers key operational aspects: state transition, confirmation requirement, failure modes, and refusal conditions. It also hints at the return format (sent:false/reason). However, it does not explain the idempotencyKey parameter or the exact response structure, leaving minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameters. It explains 'confirmed' (needs_confirmation without confirmed:true) and implies grantId/workspaceId via the action. However, it completely omits idempotencyKey, leaving the agent unaware of its purpose (e.g., deduplication). The description partially compensates but is incomplete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: sending the portal link and activating the grant (draft -> active). It specifies the resource (portal link) and the state transition, distinguishing it from siblings like portal_grant_create (which likely creates the draft) and portal_grant_revoke (which cancels). The phrase 'P8 outbound' adds specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context: it is confirm-gated (requires confirmed:true if the grant needs confirmation) and draft-gated (only works on draft grants). It also explains when it degrades (OP4 boundary) and refuses (revoked/expired). However, it does not explicitly name alternative tools or state 'use this instead of X', leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
portal_quote_acceptA
Nimm eine Offerte über das Portal an (US-F02.3): fence-checks the token (the quote must be in scope AND belong to the grant's contact) then DELEGATES to C02's real accept (the A10 sent -> accepted transition) with actor portal:. F02 owns no quote state and hand-rolls no invoice. Idempotent via idempotencyKey: a same-key re-accept returns the original, never a second acceptance (no double-issue). C02's own errors (invalid_transition on a non-sent quote, quote_expired) surface unchanged. A refused accept writes zero rows. Token-authenticated, pre-workspace: no workspaceId.
| Name | Required | Description | Default |
|---|---|---|---|
| token | Yes | ||
| quoteId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses extensive behavior: fence-checks, delegation to C02, idempotency via key, error propagation, zero-row writes on refusal, and authentication model. This far exceeds typical transparency and leaves no ambiguity about side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and technical, front-loaded with the core purpose, then elaborates on delegation, idempotency, and errors. Every sentence adds value, though it is longer than typical. The structure is logical and information-dense without being redundant, earning a 4 for efficiency despite length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description covers essential behavioral aspects: delegation mechanism, idempotency behavior, error propagation, and workspace scope. It omits explicit return format, but the idempotency description implies success behavior. For a complex operation, this is exceptionally complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains idempotencyKey in detail (same-key re-accept returns original). It implies quoteId is the quote to accept and token is the authentication token, but does not provide explicit format or validation details. While not exhaustive, it adds meaningful semantics beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'accept' and the resource 'quote' via the portal, with a use-case ID. It distinguishes itself from sibling tools like quotes_accept by explicitly noting it delegates to C02's real accept and is token-authenticated pre-workspace. The purpose is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use it (accepting a quote via portal) and provides context such as token checks, idempotency, and delegation. It does not explicitly name alternatives or state when not to use it, but the pre-workspace, token-authenticated nature implicitly differentiates it from workspace-scoped accept tools. Clear usage context is present, though exclusions are not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
portal_resolveA
Löse einen Portal-Token auf (the single verb the hosted page calls; agents call it to test a grant end to end): hashes the presented token, loads the grant, enforces expiry/revocation and the THREE FENCES (workspace -> contact -> scope), and returns ONLY the scoped read model: invoices (number, dates, total + open amount from A11/A14, QR reference read verbatim), quotes (C02 state), documents (E00 metadata). An expired or revoked token, or one scoped to another contact, returns grant_denied indistinguishably (no oracle). Token-authenticated, pre-workspace: no workspaceId. Computes no money, posts nothing; every resolve/denial lands in audit_log.
| Name | Required | Description | Default |
|---|---|---|---|
| token | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: it discloses hashing, expiry/revocation enforcement, the three-fence check, indistinguishable grant_denied behavior to avoid an oracle, that it computes no money and posts nothing, and that every resolve/denial is audit-logged. This is far beyond the schema and provides substantial behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and long, but nearly every clause adds needed behavioral or return-model detail. The core purpose is front-loaded, and the extra clauses about fences, denial indistinguishability, and audit logging earn their place. It could be split into clearer sentences, which keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema or annotations, so the description must explain both call behavior and return semantics. It does: it lists exactly what is returned for invoices, quotes, and documents; explains denial semantics for expired/revoked/other-contact tokens; clarifies auth context; and notes side effects. An agent has enough context to invoke it correctly and interpret its results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the schema only says 'token' is a string. The description compensates by explaining that the token is presented by the hosted page, is hashed for lookup, is authenticated pre-workspace, and requires no workspaceId. It does not specify token format or example values, but for a single-token parameter it adds meaningful semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear operation and resource: it resolves a portal token, enumerates the resolution pipeline (hash, load grant, enforce fences) and the exact read model returned. It also distinguishes its role from likely siblings by saying it is the single verb the hosted page calls and that agents call it to test a grant end to end. This makes it identifiable even without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: it is the hosted page's verb and used by agents to test a grant end to end, including the pre-workspace token-auth model. It does not name explicit alternatives or say when not to use it, so it stops short of a full when/when-not comparison.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_sendA
Sende die Bestellung an den Lieferanten (draft -> sent, P8 draft-gated): renders the outbound PO PDF ARTIFACT and STOPS. It returns { artifactRef, transmitted:false } and NEVER emails the supplier on its own: transmission needs the approval dial (Pattern P8/OP4). A zero-line PO cannot leave draft (no_lines); a non-draft PO is refused (invalid_transition).
| Name | Required | Description | Default |
|---|---|---|---|
| poId | Yes | ||
| actor | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It openly states that the tool does not email the supplier, that it returns a specific object with transmitted:false, and that it enforces state gating with specific error conditions. This is a very transparent description of the tool's side effects and limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a bit long but every sentence adds important information: the state transition, the artifact rendering, the return value, the non-transmission, and the error conditions. It front-loads the core behavior and then adds constraints. It is structured with a colon and clear clauses, making it scannable despite its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential behavioral context: what it does, what it returns, when it fails, and what it does not do. However, it does not explain the parameters at all, and external references like 'Pattern P8/OP4' and 'approval dial' are mentioned without elaboration, which might confuse an agent unfamiliar with the domain. Given the complexity, it is mostly complete but has minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, and the tool description does not explain any of the parameters (workspaceId, poId, actor, idempotencyKey). The agent receives no guidance on what these parameters mean or how to fill them beyond their names. This is a significant gap, as the description must compensate for low schema coverage but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool's action: rendering the outbound PO PDF artifact and stopping, not actually transmitting. It distinguishes itself from a full 'send' operation by explicitly stating it returns { artifactRef, transmitted:false } and never emails the supplier. This differentiates it from sibling tools like po_cancel or po_revise, which have different state changes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit usage guidance: it is a draft-gated step that prepares the artifact for transmission but requires a separate approval dial to actually send. It also states when the tool cannot be used: a zero-line PO cannot leave draft, and a non-draft PO is refused. This clearly tells the agent when to invoke it and when not to.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
post_entryA
Post a balanced double-entry journal entry (business sources only; reversal/close are separate tools). CONSEQUENCE: Posts an immutable journal entry; the only correction afterwards is a reversing entry.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | ||
| date | Yes | ||
| lines | Yes | ||
| source | Yes | ||
| entryId | No | ||
| description | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does well by explicitly warning that the journal entry is immutable and that the only correction path is a reversing entry. It does not mention validation failures, idempotency behavior, or permission requirements, but the immutability consequence is the most important operational trait and is clearly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The core purpose is front-loaded, and the most important consequence is bolded with 'CONSEQUENCE', making the critical information immediate and scannable for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter mutation tool with no output schema and no annotations, the description covers the core operation and its most serious consequence, but it omits return behavior, idempotency semantics, and the meaning of key parameters. This is enough for a confident tool selection but not fully sufficient for invocation without additional schema enrichment.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the purpose of most parameters such as idempotencyKey, source, ref, tax fields, or costCenter. It only hints at line balancing through 'balanced double-entry', which is relevant to the lines array but does not make up for the complete absence of parameter documentation across eight parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Post a balanced double-entry journal entry') and adds an explicit scope restriction ('business sources only'). It also distinguishes itself from sibling tools by noting that reversal and close are separate tools, which lets an agent tell post_entry apart from reverse_entry or close_period without inspecting schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this is for posting business-source journal entries, not for reversal or closing, which are separate tools. It does not exhaustively enumerate when to choose post_entry over save_draft or other entry-related tools, but the business-only qualifier and explicit separation of reversal/close provide enough practical guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
post_fx_revaluationA
Post the period-end UNREALISED currency gain/loss as a balanced entry (source fx) plus its next-period reversal, atomically. The unrealised difference books to account 6949 against each revalued position, and the entry auto-reverses on the first day of the next period so the REALISED figure at settlement (A14/A18) is never double-counted. Idempotent per periodEnd: a re-post with the same key replays, a different key returns already_posted, a zero-diff period posts nothing. Refuses needs_rate when any open FC position lacks a closing rate, and period_locked for a locked target period (A03). CONSEQUENCE: Posts the period-end unrealised currency gain or loss, with its automatic next-period reversal.
| Name | Required | Description | Default |
|---|---|---|---|
| periodEnd | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: atomicity, auto-reversal on next period, idempotency behavior per periodEnd, zero-diff no-op, and error cases. This is unusually transparent for a posting tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and front-loaded, but the final 'CONSEQUENCE' sentence redundantly restates the first sentence. Removing it would improve conciseness without losing content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema or annotations, the description covers purpose, mechanics, idempotency, and error conditions. An agent has enough context to call it correctly, including what it refuses and what it returns in edge cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters. It explains periodEnd as the period-end date, idempotencyKey semantics (replay vs already_posted), and workspaceId appears obvious from name. It does not explicitly define workspaceId but that is low-risk.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Post), resource (period-end unrealised currency gain/loss), and mechanism (balanced entry with source fx and auto-reversal). It distinguishes from sibling fx_revaluation by emphasising the posting and idempotency behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: period-end, unrealised, auto-reversal, idempotency key semantics, and refusal conditions (needs_rate, period_locked). However, it does not explicitly name sibling alternatives such as fx_revaluation_reverse or post_entry, so the exclusion is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
post_vendor_billA
Post an existing draft bill to the ledger (bucht die Kreditorenrechnung). The figures are recomputed from the stored input at the supply date, so a draft that sat across a rate change or a VAT-method change books what is correct now rather than what a stale preview cached. Refuses already_posted with the entry the bill already carries, and period_locked when the bill date falls in a locked period. CONSEQUENCE: Posts the draft vendor bill to the ledger; the only correction afterwards is a reversing entry.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| vendorBillId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses recalculation behavior, specific refusal conditions, and the irreversible consequence that only a reversing entry can correct. The labeled 'CONSEQUENCE' makes the side effect explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: action, recalculation nuance, refusal conditions, and consequence. The 'CONSEQUENCE:' label provides strong structure and front-loads the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is highly informative for selection and effect awareness, but it leaves parameter semantics—especially idempotencyKey—undocumented. With no output schema or annotations, a bit more guidance on the idempotency contract would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain workspaceId, vendorBillId, or idempotencyKey. Parameter names are somewhat self-explanatory, but idempotencyKey semantics are undocumented and additionalProperties is allowed without guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Post an existing draft bill to the ledger', reinforced by the German 'bucht die Kreditorenrechnung'. This clearly distinguishes it from draft creation, bill listing, or voiding.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context is given: it operates on an existing draft, recomputes figures at the supply date, and refuses already_posted or period_locked bills. It does not explicitly name alternatives or when not to use it, but the 'existing draft bill' framing implies the appropriate stage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_upsertA
Lege eine Bestellung an oder bearbeite sie: a purchase order (its OWN D02 document, reusing Pattern P7 on purchase_order.status, NOT an A10 document kind) on a C00 supplier contact with D00 item lines. Each item line is priced once via resolveSupplierPrice (latest supplier valid_from, else the item cost price) unless an explicit unitPriceRappen is given, and its expected tax code resolved once through A05 (H-VAT-TRACE), then snapshotted on the po_line. POSTS NOTHING and touches no stock: the financial effect stays on the A17 vendor bill (P3). A EUR order snapshots the CHF/txn fx_rate (H-FX). Passing poId edits an existing DRAFT (a sent PO is refused with invalid_transition; changes go through po_revise). A line may carry projectId, a B00 project tag for the B03 Projekterfolg (reporting only: it prices nothing; B03 reads it for accrued/committed purchase cost). A supplier that is not a C00 contact, or a line with an unknown itemId or projectId, is refused (invalid_reference); qty <= 0 is invalid_qty.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| poId | No | ||
| actor | No | ||
| lines | No | ||
| currency | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No | ||
| supplierContactId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and excels: it details pricing via resolveSupplierPrice, tax code resolution via A05, FX snapshotting for EUR, refusal conditions (invalid_reference, invalid_qty, invalid_transition), and explicit side-effect disclosure ('POSTS NOTHING and touches no stock'). This goes well beyond any annotation could.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place. It is front-loaded with the core purpose, followed by pricing behavior, financial implications, edit semantics, and validation rules. No redundant filler; the structure mirrors an effective decision tree for the agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given complexity (8 params), no output schema, and no annotations, the description covers the key operational, financial, and validation contexts. It explains error conditionsсущественные для correct invocation. The only minor gap is that it does not describe the return value or success response, but that is not strictly required without an output schema. Overall, it gives an agent enough to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the critical parameters: unitPriceRappen (explicit pricing overrides resolveSupplierPrice), poId (edit vs create), supplierContactId (must be C00), lines (pricing/tax/projectId semantics), and currency (implicitly via EUR). However, it does not describe workspaceId, actor, note, or idempotencyKey, though these are either self-evident or standard.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the action (create or update a purchase order), identifies the specific resource (a D02 document, not A10), and differentiates from siblings like po_revise ('Passing poId edits an existing DRAFT ... changes go through po_revise'). The verb and scope are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use and when-not-to-use guidance: creating or updating a draft PO vs. using po_revise for sent POs. It also clarifies that it does not post or touch stock, directing the agent to the correct alternative for financial effects. This is textbook usage differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_version_diffBRead-only
Vergleiche zwei Versionen derselben Bestellung (OP14, P5): the exact header- and line-level changes (qty, price, description, added, removed, tax) between two versions, in a structured form for GUI and agents. Pure. Two versions of DIFFERENT purchase orders are refused (invalid_reference); an unknown version id is not_found.
| Name | Required | Description | Default |
|---|---|---|---|
| toVersionId | Yes | ||
| workspaceId | Yes | ||
| fromVersionId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations include readOnlyHint=true, so the description does not need to restate read-only, but it adds 'Pure.' which reinforces that. More importantly, it discloses error behaviors (refusing different purchase orders with invalid_reference and unknown version ids with not_found), which is valuable behavioral context beyond the annotations. It also lists the types of changes captured, giving agents a preview of the output focus.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief and front-loads the core purpose. However, it mixes German and English and includes the ambiguous word 'Pure.', which adds minor noise. Overall it is reasonably concise and structured, but could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description should describe the return structure. It mentions 'structured form' and lists the fields (qty, price, description, added, removed, tax), but does not specify the exact JSON shape, nesting, or how changes are represented (e.g., old/new values). It also does not explain how version IDs are obtained or what workspaceId is for. The error cases are covered, but key operational details are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate by explaining parameters. It only mentions 'two versions' without explicitly mapping fromVersionId, toVersionId, and workspaceId. There is no explanation of parameter formats, required relationships, or how to obtain version IDs. This is a significant gap given the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool compares two versions of the same purchase order and lists the specific change dimensions (qty, price, description, added, removed, tax). It specifies the resource and action, but does not explicitly name sibling tools like po_version_get, so it does not fully distinguish from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for comparing versions, but does not explicitly state when to use it over related tools (e.g., po_version_get, po_version_list). It does mention error conditions (invalid_reference for different POs, not_found for unknown IDs), which helps agents decide when the tool is applicable, but there is no explicit guidance on alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_version_getBRead-only
Lies eine einzelne, eingefrorene PO-Version (OP14, P5): the full immutable snapshot (header + lines) of ONE version, exactly as it stood when that version became active. The Treuhänder-Rekonstruktion of the commitment at a point in time (OR 957a). A version id from another workspace is refused (not_found, §H-TENANT).
| Name | Required | Description | Default |
|---|---|---|---|
| versionId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, which the description aligns with. The description adds the immutable snapshot behavior and the cross-workspace rejection (not_found). However, it does not mention potential rate limits, authentication requirements, or what happens on non-existent versions beyond the cross-workspace case. It adds some value but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core purpose. It packs a lot of information into three sentences with no fluff. The only minor issue is the mixed-language phrasing and the parenthetical §H-TENANT, which might be less clear to a general agent, but overall it is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only tool with two parameters, the description covers the core behavior. However, given the lack of output schema and parameter descriptions, it could benefit from explaining the return structure or stating that versionId is opaque. The cross-workspace behavior is covered, but other error cases are not. It is adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the schema itself provides no descriptions for the two parameters. The description indirectly refers to workspaceId via the cross-workspace note but does not provide explicit syntax or format for versionId or workspaceId. It adds a little meaning but leaves the agent to infer parameter purposes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (read a single frozen PO version) and the resource (PO version snapshot). It distinguishes from siblings by emphasizing immutability and the full snapshot. However, it does not explicitly name a sibling like po_get or po_version_list, so differentiation is implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (to reconstruct commitment at a point in time, matching OR 957a) but does not explicitly state when not to use it or provide alternatives. It hints at the cross-workspace refusal, but usage guidance is more implicit than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
po_version_listBRead-only
Liste die Versionshistorie einer Bestellung (OP14, P5): the ordered version trail (1..N) of a purchase order, each row carrying its number, status (active | superseded), the amendment reason that produced it, actor, timestamp and the frozen sent artifact. Materialises version 1 on first touch, so a legacy D02 PO created before I01 shows a complete history. An empty PO id is refused (not_found).
| Name | Required | Description | Default |
|---|---|---|---|
| poId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotation Contradiction: annotations declare readOnlyHint=true, but the description states the tool 'Materialises version 1 on first touch', which is an explicit side effect that creates a version row. Even though the description is otherwise transparent about the materialisation and error behavior, this direct conflict with the read-only annotation forces a score of 1.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with purpose, itemizes the row fields, and adds materialisation and error behaviour in a compact paragraph. The bilingual repetition and domain-specific codes (OP14, D02, I01) add mild noise, but every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list operation with two obvious parameters, the description covers the returned row fields, the lazy materialisation effect, and the empty-id refusal. It lacks an explicit return contract and workspaceId semantics, but these are minor gaps given the description's detail and the absence of an output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only adds that an empty PO id is refused and does not explain the semantics or format of poId or workspaceId. The general reference to a purchase order gives some context, but the description does not map clearly to both required parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description leads with a concrete verb and resource: 'Liste die Versionshistorie einer Bestellung' and clarifies it as 'the ordered version trail (1..N) of a purchase order'. This clearly defines a list-versions operation and differentiates it from siblings like po_version_get or po_version_diff by describing the full trail and row contents.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when it applies (listing all versions of a PO, including legacy POs) and gives a useful constraint (empty PO id refused). However, it does not explicitly mention alternatives or state when to use po_version_get/diff instead, leaving some selection logic to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_feedbackA
Write a feedback report to ~/.till/feedback and return a mailto link for it. TILL cannot send: the report is handed to the operator to send from their own mail client.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| locale | No | ||
| message | Yes | ||
| subject | Yes | ||
| clientError | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| includeDiagnostics | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full transparency burden and discloses the key behaviors: it writes a file to a specific local path, does NOT send anything, and returns a mailto link for manual sending. It leaves idempotency behavior and overwrite/append semantics undisclosed, which matter given the required idempotencyKey parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences of about 33 words, with the core action front-loaded (write to path + return mailto link) followed by the critical boundary constraint. Both sentences each carry load-bearing information with no filler. This is efficiently structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters (0% documented), a nested clientError object, no output schema, no annotations, and a mailto return shape, the description covers the core behavior but leaves parameter semantics, idempotency semantics, and the mailto link structure unexplained. It covers the essentials for a first correct call, making it minimally viable but with clear gaps for a tool this complex without structured support.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% — no parameter has any description — so the tool description must compensate. It only implicitly covers subject and message; kind, locale, clientError, idempotencyKey, and includeDiagnostics remain unexplained, and the structure of the returned mailto link is unspecified. An agent would be guessing on most inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Write) with a precise resource and destination (~/.till/feedback) and declares the return value (mailto link), making the tool's purpose unmistakable. The write-and-return-mailto action clearly distinguishes it from sibling tools like preview_feedback and list_feedback even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear workflow context: TILL cannot send email, so the report is generated by this tool and handed to the operator for out-of-band sending. That tells an agent when this tool is the right choice versus a sending path, but it stops short of naming alternatives or stating explicit when-not-to-use conditions, so not a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
prepare_periodA
Ready a period for the Treuhänder's review: builds one packet (review coverage, draft count, unmatched bank items from the camt and QR queues, open debtors, MWST preview) and leaves machine flags on anomalies it detects (a duplicate-looking posting, a line missing a tax code where the account declares a default). It never approves, locks, or exports: those stay human acts. Idempotent per period, a re-run refreshes the packet and never duplicates a flag.
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and excels: it discloses that the tool leaves machine flags on anomalies, gives concrete anomaly examples, and states the idempotency guarantee that re-runs refresh the packet without duplicating flags. It also draws a clear boundary around what it will never do (approve, lock, or export).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences carry substantial information with no filler, front-loading the purpose before adding packet contents, boundary conditions, and idempotency. Every clause earns its place and the structure is easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is complete enough for a side-effect preparation tool: it defines inputs' role, outputs' purpose, anomaly behavior, and safety. It omits explicit return-value semantics and parameter format constraints, which matters slightly because no output schema is available, but the core decision context is solid.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only bare string types with zero coverage, but the description clarifies the meaning of period and idempotencyKey through 'per period' and the re-run/no-duplicate-flag behavior. WorkspaceId is not described, though its role is strongly inferable from the name and the standard workspace-based API context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening phrase 'Ready a period for the Treuhänder's review' states a clear verb, resource, and audience, and the description enumerates what the packet contains. It also explicitly distances itself from approve/lock/export operations, making it easy to distinguish from sibling tools like approve_entry, lock_period, and export_journal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear workflow context: call it to build review material for the Treuhänder, and it says re-runs are safe because the operation is idempotent. It explicitly says finalizing actions are excluded ('those stay human acts'), but it does not name sibling alternatives, so routing guidance is contextual rather than fully explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_bank_opening_balanceARead-only
Show what set_bank_opening_balance WOULD post for a Bankkonto, in the workspace base currency, without posting anything: the base amount, the rate it converted on, and the two legs (the Bankkonto's verknüpftes Konto against 9100 Eröffnungsbilanz). It runs the posting's own arithmetic, so the FIGURE cannot disagree, and it answers the same refusals the posting answers for the state it reads (needs_account for a missing or archived 9100, currency_mismatch, opening_balance_already_set, needs_fx_rate, period_locked). It is not a promise about the posting. It checks no capability, so a caller without the post right gets the full figure here, and at set_bank_opening_balance that caller is refused permission_denied at every amount, including the zero one that books no journal entry; that refusal is not always the FIRST one, because an unknown Bankkonto, an opening balance already set, or (above zero only) a missing 9100 is answered before the capability is consulted, and an idempotencyKey that replays an already-completed call still replays. It also holds no state, so a period locked, a 9100 archived or another actor posting in between still refuses there. Takes no idempotency key and can never post: use it before Buchen, because a posted opening balance is immutable and its only correction is a reversing entry.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | ||
| fxRate | No | ||
| currency | No | ||
| amountMinor | Yes | ||
| workspaceId | Yes | ||
| bankAccountId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint:true, but the description discloses much more: the tool never posts, holds no state, checks no capability, is not a promise about the posting, and mirrors specific refusal conditions such as needs_account, currency_mismatch, and period_locked. No contradiction with the read-only annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long and dense, but every sentence contributes distinct operational facts: posting dry-run, refusal parity, no capability check, no state, no idempotency key, immutability of posted balances. It is front-loaded with the core purpose and the later details are earned given the complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a preview tool with no output schema, the description states exactly what the preview returns (base amount, rate, two legs), which error conditions it mirrors, and the critical caveat that it is not a promise about the posting. An agent has sufficient operational and safety context to invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not systematically explain any parameter. It provides loose context by mentioning the base amount, conversion rate, and Bankkonto, but it never clarifies workspaceId, bankAccountId, amountMinor as minor units, date, currency, or fxRate. Parameter names are mostly self-explanatory, so this is not a 1, but the gap is real.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb ('Show') and clearly identifies the resource and operation ('what set_bank_opening_balance WOULD post for a Bankkonto'), while immediately stating it posts nothing. This distinguishes it decisively from the sibling set_bank_opening_balance and leaves no ambiguity about its role as a dry-run preview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly instructs 'use it before Buchen' and explains why: a posted opening balance is immutable and only correctable by a reversing entry. It also describes the preview's divergence from the real posting (no capability check, no state held, same refusals), giving an agent clear when-to-use and expectations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_document_templateARead-only
Render a template preview as PDF bytes, mutating nothing: against sampleDocumentId, else the most recent real document of the template's kind, else synthetic MUSTER-watermarked sample data that is never a payable document (no fabricated QR code; a missing QR-IBAN surfaces A11's own needs_qr_iban cause).
| Name | Required | Description | Default |
|---|---|---|---|
| templateId | Yes | ||
| workspaceId | Yes | ||
| sampleDocumentId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the readOnlyHint annotation by disclosing fallback precedence, synthetic MUSTER-watermarked data, that it is never a payable document, that it does not fabricate a QR code, and that a missing QR-IBAN surfaces A11's own needs_qr_iban cause. This is rich behavioral context with no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the action and safety guarantee, then adds the fallback chain and watermark/QR caveats. Every clause carries distinct, decision-relevant information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter read-only preview tool, this is complete: it states the output (PDF bytes), the non-mutating guarantee, the data-source fallback chain, and the important sample-data caveats. No output schema exists, but the description already covers the essential return and behavior information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by explaining the optional sampleDocumentId as the preferred preview source and implying templateId through 'the template's kind'. It does not describe workspaceId or parameter types, so it is helpful but not exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with 'Render a template preview as PDF bytes, mutating nothing', naming the verb, resource, output format, and non-mutating nature. It clearly distinguishes this from list/get/update/archive template siblings by framing it as a preview action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage context: use it to render a preview against sampleDocumentId, otherwise fall back to the most recent real document of the template's kind, otherwise synthetic sample data. It does not explicitly name sibling alternatives or state when not to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_feedbackARead-only
Render exactly what a feedback report would contain, writing nothing. Use it to show a person what would be sent before it is.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| locale | No | ||
| message | Yes | ||
| subject | Yes | ||
| clientError | No | ||
| workspaceId | Yes | ||
| includeDiagnostics | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description reinforces the readOnlyHint with 'writing nothing' and adds the fidelity promise of rendering exactly what would be sent. No contradiction exists, but the safety behavior is already carried by annotations and no further behavioral detail is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short, front-loaded sentences with no filler. Every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The high-level purpose is clear, but 7 parameters are unexplained and there is no output schema to fill in the gaps. The tool is not fully callable by an agent without guessing at parameter semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage and 7 parameters, the description gives no meaning for workspaceId, subject, message, kind, locale, clientError, or includeDiagnostics. The agent is left without parameter-level guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Render exactly what a feedback report would contain'. 'Writing nothing' clearly differentiates it from mutating or send-like siblings such as prepare_feedback or list_feedback.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use it to show a person what would be sent before it is' gives clear timing and context: invoke it before sending a feedback report. It does not explicitly name alternatives or exclude cases, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_opening_importARead-only
Dry-run already-parsed migration rows into an opening position without writing anything: normalises a signed balance column or a debit/credit pair into integer Rappen, resolves each account number against the chart, and reports the rows the chart does not know (unmapped) alongside the balance check. Takes rows as an array of column-value objects, NOT a CSV string: parsing the export into rows is the caller's step. It refuses the same things import_opening_balances refuses, including an account named twice, so a preview that comes back clean is not followed by a surprise rejection. This reports what the FILE says and whether its own two sides agree; it cannot tell you the file is the right one, so compare the per-account lines against the Beleg. It takes no idempotency key and can never post.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| rows | Yes | ||
| format | No | ||
| mapping | No | ||
| reference | No | ||
| workspaceId | Yes | ||
| differenceAccount | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations declare readOnlyHint=true, but the description goes further by stating 'without writing anything' and 'can never post.' It also discloses the normalization process, account resolution, reporting of unmapped rows, refusal of duplicate accounts, and the inability to verify file correctness—rich context beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence contributes: purpose, input format, validation behavior, limitation, and idempotency. It is front-loaded with the dry-run nature and avoids fluff, though it is slightly long.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 7 parameters, a nested mapping object, and no output schema, the description covers the core behavior and input requirements, including what it reports (unmapped rows and balance check). It omits details on optional parameters and the exact output structure, but these are secondary to the tool's main use case.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate. It explains the rows parameter (array of column-value objects, not CSV) and touches on balance/debit/credit normalization, but does not explain asOf, format, reference, differenceAccount, or the mapping object's field semantics. It adds some value but leaves significant gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: a dry-run of migration rows into an opening position, with specific actions like normalizing balances and resolving account numbers. It distinguishes itself from import_opening_balances by being read-only and taking parsed rows, making it easy to differentiate from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names import_opening_balances as the sibling it mirrors and clarifies the input format (rows array, not CSV). It also provides a limitation (cannot tell if the file is correct) and advises comparing against the Beleg, giving clear when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_open_itemsARead-only
Preview an open-items migration for one plan and side (ar Debitoren or ap Kreditoren): compute the control tie-out the batch WOULD produce (the migrated open total against the opening 1100 Debitoren or 2000 Kreditoren line, integer Rappen, zero tolerance, G11 three-status honesty) and surface every refusal (open_item_total_mismatch, contact_unmapped/vendor_unmapped, tax_unresolved, needs_fx_rate) naming the offending row, WITHOUT writing anything. priorYearDetail (live or archive, default archive) governs whether an already-settled prior-year row is carried live or belongs in the G13 archive.
| Name | Required | Description | Default |
|---|---|---|---|
| rows | Yes | ||
| side | Yes | ||
| planId | Yes | ||
| workspaceId | Yes | ||
| priorYearDetail | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the readOnlyHint annotation by stating 'WITHOUT writing anything' and detailing the exact refusal types it surfaces, the zero tolerance, integer Rappen, and G11 three-status honesty. It also explains the priorYearDetail option's effect. This is rich behavioral disclosure that aligns with and extends the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but dense with essential information. It front-loads the purpose and key constraints, then lists refusal types and the priorYearDetail behavior. Every sentence contributes value; there is no filler. It is structured logically, moving from what it does to how it behaves to the configurable option.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (5 params, no output schema), the description covers the main behavior, refusal types, and the priorYearDetail switch. It does not specify the exact return structure, but it names the refusal categories and indicates it surfaces them with offending rows, which is sufficient for an agent to understand the outcome. It is reasonably complete for a preview tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by explaining the key parameters: planId and side are implied by 'one plan and side', rows are referenced as the input rows that may cause refusals, and priorYearDetail is explicitly described. workspaceId is not detailed, but the description provides enough semantic context for an agent to understand the tool's purpose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool previews an open-items migration for a specific plan and side, computes a control tie-out, and surfaces refusals without writing. It distinguishes itself from migration execution tools by explicitly stating it is a preview and names the exact refusal types. The verb 'preview' is specific, and the resource is open-items migration.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is a dry-run before committing a migration, saying it computes what the batch WOULD produce and surfaces refusals without writing. It does not explicitly name alternative tools, but the preview nature and the mention of priorYearDetail behavior give clear context for when to use it. There is no explicit when-not-to-use, but the intent is unmistakable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_paymentARead-only
Preview exactly what a payment would post, without writing anything: the remainder still to allocate, each document's resulting open amount and status, the Skonto VAT split, the Ist-timing paid-portion VAT, the currency conversion with its realised difference, and the journal legs with account numbers and labels. Call this before record_payment so the figures a human sees and the figures the ledger books are the same figures. It takes no intent and can never post.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | ||
| fxRate | No | ||
| source | No | ||
| currency | No | ||
| direction | Yes | ||
| reference | No | ||
| allocations | No | ||
| amountMinor | Yes | ||
| workspaceId | Yes | ||
| bankAccountId | Yes | ||
| counterpartyId | No | ||
| onAccountMinor | No | ||
| counterpartyKind | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description reinforces this with 'without writing anything' and 'can never post,' which adds helpful safety context. It also enumerates the detailed outputs it returns, giving the agent a clear picture of what the preview will contain.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the key point and enumerates outputs in a dense but organized way. Some redundancy exists in phrases like 'without writing anything' and 'can never post,' but overall the description is appropriately structured for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 13 parameters, no output schema, and no parameter descriptions, the description provides useful output context but leaves input semantics largely unaddressed. It explains what the preview produces and when to call it, making it minimally viable, but not complete enough for an agent to confidently construct all parameter values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 13 parameters, so the description carries a heavy burden to explain how inputs relate to concepts like allocations, skonto, VAT splits, and currency conversion. It names some concepts in outputs but does not map them to parameters or clarify semantics for required fields. This is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the operation is a read-only preview of what a payment would post and explicitly names record_payment as the mutating counterpart. This strongly differentiates it from sibling payment tools like record_payment and allocate_payment.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives direct usage guidance: 'Call this before record_payment' and emphasizes that it 'can never post' and takes no intent. This tells the agent exactly when to use it and distinguishes it from the related mutating payment tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
preview_plugin_installARead-only
Prüfe eine Erweiterung vor der Installation (US-G02.1): parse a .tillplugin bundle`s manifest and return its name, version, declared capabilities (which MCP tools, Studio screens, report sources, automation actions it registers) and requested permission scopes, plus whether it is compatible with the current core, WITHOUT persisting anything and WITHOUT starting a sandbox. A malformed manifest answers invalid_manifest naming what failed to parse.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | ||
| packageRef | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds meaningful behavioral context beyond that: parsing the manifest, persisting nothing, starting no sandbox, and answering invalid_manifest for malformed bundles with the parse failure named. This is genuinely useful additional context and does not contradict anything in the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is effectively front-loaded and the capability-category list is informative, but the text is one dense run-on sentence mixing German and English with a stray requirement ID (US-G02.1) and a typo ('.tillplugin' bundle's'). Splitting it into shorter sentences and trimming the noise would improve scanability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema present, the description compensates reasonably by spelling out the returned metadata fields and the invalid_manifest error case. The major gap is the complete absence of input-parameter guidance plus no explicit statement of the success signal, leaving an agent uncertain how to construct the three required arguments.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for all three required parameters (workspaceId, source, packageRef), and the description never explains what each parameter carries or how they reference the bundle. The mention of '.tillplugin bundle' lets an agent infer that source/packageRef relate to the plugin package, but the required input semantics remain essentially guesswork.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Prüfe eine Erweiterung vor der Installation' — check an extension before install) and precisely enumerates what is returned: name, version, declared capabilities, permission scopes, and core compatibility. This clearly differentiates it from sibling tools like install_plugin, enable_plugin, and the registry lookup tools even before examining their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'vor der Installation' frames this as the pre-install check step, and the explicit negations ('WITHOUT persisting anything and WITHOUT starting a sandbox') further clarify it is a safe inspection to run before invoking install_plugin. However, it never names a sibling alternative or states an exclusion like 'call install_plugin only after this returns compatible', so it just misses the explicit when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
price_lists_deleteA
Delete a price list together with its price rows. A price row has no reader apart from its list and a document line already snapshots the price it resolved, so the cascade removes nothing any issued document depends on. Refused with price_list_referenced if anything still points at the list.
| Name | Required | Description | Default |
|---|---|---|---|
| priceListId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does well: it reveals the cascade behavior, explains why the cascade is safe by noting document lines snapshot resolved prices, and discloses the specific error code price_list_referenced. It does not mention irreversibility or permission requirements, but the main side-effect profile is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly written sentences, each earning its place: the first states the action, the second justifies the cascade safety, and the third gives the failure condition. No filler, and the key purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no annotations and no output schema, the description covers the essential operational facts: what gets deleted, why downstream documents are unaffected, and when the operation is refused. It omits only minor details like response format or behavior on non-existent IDs, which are not critical for invoking this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needed to compensate for undocumented parameters, but it does not. It never mentions priceListId, workspaceId, or idempotencyKey, nor their roles. The name 'price list' loosely maps to priceListId, but the description adds no meaning beyond the raw schema property names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Delete a price list together with its price rows.' It clearly distinguishes this destructive operation from sibling price-list tools like price_lists_upsert, price_lists_set_price, and price_lists_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives useful operational context: it explains when deletion is safe (nothing issued depends on the price rows) and when it will fail ('Refused with price_list_referenced if anything still points at the list'). It does not name alternative tools explicitly, but the precondition and refusal condition provide clear usage guidance for a destructive operation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
price_lists_getBRead-only
Read one price list and its price rows (the validFrom history per item).
| Name | Required | Description | Default |
|---|---|---|---|
| priceListId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes that this is a safe read operation, and the description agrees. It adds some output context by mentioning price rows and the validFrom history, but it does not disclose not-found behavior, pagination, or whether archived/deleted price lists are readable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It states the operation and the scope of returned data efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read operation with a readOnly hint and no output schema, the description is minimally viable: it identifies the resource and the key returned data. However, it leaves parameter semantics undocumented and does not describe the response shape in enough detail, so the agent must infer several operational details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no explanation of workspaceId or priceListId beyond the vague reference to 'one price list.' It does not compensate for the schema's lack of parameter documentation, leaving the agent to infer parameter meanings from names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Read') and a concrete resource ('one price list'), then clarifies what is returned: price rows and the validFrom history per item. This distinguishes the tool from price_lists_list and from mutation siblings like price_lists_upsert or price_lists_set_price.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Read one price list' implies a single-record retrieval use case, but the description never explicitly states when to use this tool instead of price_lists_list or mentions any alternatives/exclusions. Usage guidance is inferred from the singular scope and naming rather than clearly articulated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
price_lists_listBRead-only
List the workspace price lists with their scope (contact or segment).
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotation readOnlyHint=true already signals that the operation is safe and non-destructive. The description adds the detail that the output includes the scope (contact or segment), which is useful but does not go beyond that. It does not disclose any pagination, ordering, or potential size limits, but given the read-only annotation, a baseline of 3 is acceptable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence with no fluff. It front-loads the action and resource, then adds the scope detail. Every word adds value, and there is no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a simple list tool with one parameter and no output schema, which lowers the complexity bar. However, the description fails to explain the workspaceId parameter, and it does not hint at the return format or any list-level behavior (e.g., ordering, limits). For a tool that returns data, the lack of return type description is a notable gap, especially since schema coverage is zero.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not mention the workspaceId parameter at all. The agent has no textual explanation of what workspaceId means, its format, or any constraints. Since the schema provides only a type (string) and no further guidance, the description fails completely to compensate for the missing schema doc.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the verb 'List' and the resource 'workspace price lists', and it includes the meaningful detail that each list carries a scope of 'contact or segment'. This clearly distinguishes it from related tools like price_lists_get, price_lists_upsert, and price_lists_set_price, and leaves no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention that it is the listing counterpart to price_lists_get or price_lists_upsert, nor does it state any conditions that would prefer this tool over others. The agent is left to infer usage from the name and sibling context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
price_lists_set_priceA
Set a list price for an item from a validFrom date (ISO YYYY-MM-DD; anything else is refused with invalid_input, never coerced). Price history is append-only: this adds a row rather than editing a past one, so a resolve at any date reads a stored integer Rappen amount.
| Name | Required | Description | Default |
|---|---|---|---|
| itemId | Yes | ||
| currency | No | ||
| validFrom | Yes | ||
| priceMinor | Yes | ||
| priceListId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses append-only behavior, strict ISO date validation with explicit failure mode, and that prices are stored as integer Rappen. It also explains the read-side implication, which is valuable for understanding side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no redundant text. The core action is front-loaded, and both sentences add essential behavioral details. No waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential behavioral contract (append-only, strict date, integer amount) but leaves several parameters undocumented in terms of semantics and does not describe the response. Given the absence of an output schema, some information is missing, but the description is still fairly complete for a simple mutation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains validFrom format and that priceMinor is an integer Rappen amount, but does not clarify workspaceId, priceListId, itemId, currency, or idempotencyKey. This is partial compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the action clearly: 'Set a list price for an item from a validFrom date.' It also distinguishes from update-like operations by explaining the append-only nature. The sibling price_lists_unset_price and price_lists_delete are clearly different actions, so an agent can tell them apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Does not explicitly state when to use this tool vs alternatives like price_lists_upsert or price_resolve. It implies usage for setting a new price from a given date, but does not mention exclusions or conditions. This leaves some ambiguity for agents comparing siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
price_lists_unset_priceA
Remove an item price from a list: pass validFrom to retract exactly that dated row, omit it to remove the item from the list entirely. Removing the item entirely is what makes price_resolve fall through to the next scope and what lets delete_item stop counting the list as a reference. Answers how many rows it took away, so a replay reports the same count and a repeat reports none.
| Name | Required | Description | Default |
|---|---|---|---|
| itemId | Yes | ||
| validFrom | No | ||
| priceListId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden and does so admirably. It discloses idempotency semantics ('a replay reports the same count and a repeat reports none'), the exact effect of parameter presence/absence, and side effects on price_resolve and delete_item.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences deliver the core action, both usage modes, side effects, and return behavior without wasted words. The conditional guidance is front-loaded, making the most important decision (pass or omit validFrom) immediately visible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is complete for a parameterized mutation tool: it explains what happens in each mode, what the return value tells the caller, and how the operation interacts with related tools. The required parameters are declared in the schema, and the description adds the operational context needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains validFrom's nuanced behavior (retract dated row vs remove item entirely), which is the most ambiguous parameter. The remaining parameter names (workspaceId, priceListId, itemId, idempotencyKey) are reasonably self-explanatory from their names and required status.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Remove an item price from a list') and clearly distinguishes the two modes of operation. It differentiates itself from siblings like price_lists_set_price and price_lists_delete by describing the exact retraction behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit guidance on when to pass validFrom versus omit it, and explains the downstream consequences of full removal. It does not explicitly name alternatives or exclusions, but the usage context is clear enough for correct invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
price_lists_upsertA
Create or edit a price list, scoped to exactly one contact OR one segment (scope_ambiguous otherwise). One scope holds at most one list: a second list for a contact or segment that already has one is refused with scope_taken naming the existing list, so no price is ever resolved by insertion order. Pass priceListId to edit, omit it to create.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| segment | No | ||
| contactId | No | ||
| priceListId | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It candidly explains the scope_taken refusal behavior VERB and the rationale ('no price is ever resolved by insertion order'), which is critical for understanding the upsert semantics. It also hints at ambiguity handling ('scope_ambiguous otherwise'), though it does not detail idempotency key behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and every sentence contributes new information: purpose, scoping constraint, uniqueness rule, and create/edit indicator. It is slightly dense but not bloated, and it front-loads the primary action before diving into edge cases.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (unique-per-scope upsert with ambiguous-scope cases) and the absence of an output schema, the description covers the most essential rules but leaves out parameter details and return behavior. It is adequate for a basic call but incomplete for nuanced scenarios like idempotency or required field validation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the role of priceListId (edit vs create) and the contact/segment scoping, but it omits the meaning of workspaceId (a required parameter), name, and idempotencyKey. This is a significant gap for a tool that agents will call with only this description as guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Create or edit') and resource ('price list'), and immediately distinguishes it from common alternatives like price_lists_set_price or price_lists_delete. The scoping constraint ('exactly one contact OR one segment') adds precision and sets expectations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides concrete guidance on when to create versus edit ('Pass priceListId to edit, omit it to create') and explains the one-list-per-scope rule, which helps users avoid common mistakes. It does not explicitly name sibling alternatives, but the create/edit framing makes the tool's niche clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
price_resolveARead-only
Resolve the effective price of an item for a contact at a date: precedence contact then segment then base (the item base sales price), the latest validFrom that is in force winning within a scope. at is an ISO YYYY-MM-DD day (or a full ISO instant) and defaults to today; any other format is refused with invalid_input rather than compared, because a date that does not sort silently resolves the wrong price. Returns priceMinor, currency, and the source.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | ||
| itemId | Yes | ||
| contactId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers the safety profile, and the description adds materially beyond it: strict ISO-format validation on `at` with an explicit invalid_input refusal, the rationale that a non-sorting date silently resolves the wrong price, the today() default, the latest-validFrom tie-break, and the return contract (priceMinor, currency, source). This is precisely the behavioral context an agent needs beyond what structured annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core function and precedence, followed by the date contract and its rationale, closing with the return shape. The rationale sentence ('because a date that does not sort silently resolves the wrong price') earns its place by teaching the agent why strict formatting matters. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and only a readOnly hint, the description carries the full burden and covers the essentials: computation rule, date handling, error behavior, and return values. Minor gaps remain — 'segment' and 'within a scope' are not precisely defined, and the relationship to price_lists_set_price/price_lists_get is unstated — but an agent has what it needs to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description must carry parameter meaning, and it does: `at` receives full treatment (ISO day or instant, defaults to today, strict validation), and itemId/contactId are mapped contextually via 'of an item for a contact.' workspaceId is left to convention, and the term 'segment' is named but not defined, so it stops short of a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb + resource: 'Resolve the effective price of an item for a contact at a date,' which clearly distinguishes this from raw price-list lookups (price_lists_get) and rate resolution (time_resolve_rate). The precedence chain 'contact then segment then base' defines exactly what computation this tool performs, so there is no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states the resolution semantics — precedence and latest-validForm selection within a scope — giving an agent clear context for when this tool is the right choice: whenever the effective, precedence-resolved price is needed rather than a raw list lookup. It does not explicitly name alternative sibling tools or state when not to use it, which keeps it from a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
procurement_anomaliesARead-only
Beschaffungs-Anomalien (procurement anomalies) since a date (default 30 days): a prioritised, DERIVED-not-stored list of unusual events with type, severity (info / warning / critical), a summary, the related PO / bill / supplier / requisition id, the detection date and a structured payload. Phase-1 types: large_price_variance, match_override_high_value, open_commitment_aging, stalled_requisition and grir_material_exposure. Thresholds are sensible defaults; this verb never writes them. Recomputed on every read, so the next read reflects the current state. Empty answers ok with an empty list. Filter by type or severity.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| since | No | ||
| types | No | ||
| severity | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description adds significant behavioral context: it is DERIVED-not-stored, recomputed on every read, never writes thresholds, and returns an empty list when there are no anomalies. This goes well beyond the annotation and provides clear expectations about side effects and freshness.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph that front-loads the core purpose, then lists types, then behavior, then filtering. Every sentence adds value—there is no fluff or repetition. It is concise yet comprehensive, achieving high information density without being verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description must explain the return shape, which it does: it lists the fields (type, severity, summary, related IDs, detection date, payload). It also explains derived/recomputed behavior and empty-list handling. However, it does not specify the payload structure in detail, nor does it mention pagination or the default/expected limit behavior. For a tool with this complexity and no output schema, a bit more detail on the response format would be ideal, so 4 is fair.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has no parameter descriptions (0% coverage), so the description carries the burden. It explains 'since' with a default of 30 days, lists the types that can be used for filtering, and mentions severity options (info/warning/critical). It does not describe the 'limit' parameter or the exact format for 'since' (e.g., ISO date), and 'workspaceId' is only implicitly required. While it covers most parameters meaningfully, it omits 'limit' and leaves some format details unspecified, so a 4 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that this tool returns a derived, prioritized list of procurement anomalies with specific fields (type, severity, summary, related IDs, detection date, payload). It lists the phase-1 anomaly types and explicitly notes it is derived-not-stored, distinguishing it from stored reporting tools. This is a specific verb+resource with clear scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions that thresholds are defaults and that this verb never writes them, implying read-only usage, and mentions filtering by type or severity. However, it does not explicitly state when to use this tool versus alternative procurement tools like procurement_open_commitments or procurement_match_status. There is no direct comparison or exclusion of alternatives, so the guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
procurement_grir_clearingARead-only
GR/IR-Abgrenzung (goods-received / invoice-received clearing) as of a date: two complementary sets from the live po_line counters, received-not-invoiced (received_qty > billed_qty) and invoiced-not-received (billed_qty > received_qty), each with its residual quantity and residual value at the PO base price (integer CHF Rappen), plus the net exposure. Status is cleared when both sides are empty within the materiality filter, else exposure_present. A pure projection that quantifies accrual exposure before period close; it never posts the accrual itself (that stays the period-close process). include_detail returns the per-line rows.
| Name | Required | Description | Default |
|---|---|---|---|
| as_of | No | ||
| workspaceId | Yes | ||
| supplier_ids | No | ||
| include_detail | No | ||
| materiality_rappen | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description reinforces this by stating it 'never posts the accrual itself'. It further discloses the status logic ('cleared' vs 'exposure_present'), the unit (CHF Rappen), the source from 'live po_line counters', and the effect of include_detail. This goes well beyond the annotation to fully explain behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause contributes. It front-loads the core definition and uses parentheses and semicolons to pack in edge cases (integer CHF Rappen, status clearing condition) without rambling. No sentence is filler; it is efficient for a tool with this complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter tool with no output schema, the description explains the logical output (two sets, residual values, net exposure, status) and the include_detail behavior. It does not specify the exact field names or formatting of the response, but since there is no output schema, the description gives enough for an agent to understand the result. Minor gap: supplier_ids semantics are not spelled out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameters. It does so for as_of ('as of a date'), materiality_rappen ('materiality filter'), and include_detail ('returns the per-line rows'). supplier_ids is not explicitly mentioned, but its purpose can be inferred from the overall clearing context. The description adds meaning beyond the bare schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly names the tool's function ('GR/IR-Abgrenzung') and defines its two complementary sets (received-not-invoiced and invoiced-not-received) with residual quantities and values. It precisely distinguishes this from posting accruals by stating it is a 'pure projection' and 'never posts the accrual itself', which differentiates it from period-close tools like accrual_create or provision_post.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: it quantifies accrual exposure 'before period close' and clarifies what it does not do ('never posts'). It does not name specific sibling tools or explicitly state when NOT to use it, but the projection/posting distinction gives sufficient guidance for an agent to choose this over alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
procurement_landed_cost_varianceARead-only
Landed-Cost-Abweichung (landed-cost variance): for each allocated I03 voucher in the window, the planned total cost against the capitalized and expensed (variance) split the confirm actually posted, with the freight / duty / insurance / other component breakdown and the variance percentage. Largest absolute variance first. All figures are integer CHF Rappen read from the live landed_cost_voucher / landed_cost_line; posts nothing. No I03 activity answers ok with an empty list.
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No | ||
| format | No | ||
| to_date | No | ||
| from_date | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description explicitly states 'posts nothing', names the live source tables (landed_cost_voucher / landed_cost_line), specifies integer CHF Rappen, states ordering by largest absolute variance, and defines the empty-result behavior. This is rich behavioral disclosure with no contradiction to the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and every clause carries useful information: scope, comparison basis, component breakdown, metric, ordering, data source, unit, read-only nature, and empty-result behavior. The opening German name is a minor duplication, but it does not bloat the description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex read-only report with no output schema, the description covers the output shape, source, number semantics, ordering, and no-data case very well. It is slightly incomplete only in that parameter semantics for filter/format are not explained, but the required workspaceId and date-window concept are at least inferable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the semantics of filter, format, from_date, to_date, or workspaceId beyond the indirect phrase 'in the window'. The property names for dates carry some meaning, but the opaque filter and format parameters are left undocumented, so the description does not adequately compensate for the missing schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly defines a specific analytical report: for each allocated I03 voucher it compares planned total cost against the actual posted capitalized/expensed split, includes component breakdown and variance percentage, and specifies ordering. This unambiguously distinguishes it from the landed_cost_create/allocate/confirm/reverse/list/get siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage context is explicit: it applies to allocated I03 vouchers in a window and is a read-only variance analysis, and the 'No I03 activity answers ok with an empty list' clause sets expectations for empty results. It does not explicitly name alternative tools, so it falls short of a 5, but the context is clear enough for an agent to select it over the landed-cost mutation tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
procurement_match_statusARead-only
Status und Ausnahmen des Belegabgleichs (three-way match status & exceptions): the summary counts and matched value by status (matched / partial / overridden) plus one exception row per non-fully-matched I04 record, with its quantity and price variance in Rappen, whether tolerance was breached, the override reason, the aging since the match and a suggested next action (await_receipt / review_override / escalate). Status is clean when there is no exception in the filter, else exceptions_present. Reads the live I04 three_way_match active rows and posts nothing; never auto-fixes a match. Filter by status, supplier or match date; include_detail expands the per-line variances.
| Name | Required | Description | Default |
|---|---|---|---|
| as_of | No | ||
| format | No | ||
| status | No | ||
| to_date | No | ||
| from_date | No | ||
| workspaceId | Yes | ||
| supplier_ids | No | ||
| include_detail | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description reinforces this by stating it 'posts nothing' and 'never auto-fixes a match.' It also discloses the output shape (summary counts, matched value by status, exception rows with quantity/price variance, tolerance breach, override reason, aging, suggested next action) and the status semantics (clean vs exceptions_present). This goes beyond the annotation by explaining what the agent can expect from the response and the operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: it front-loads the tool's purpose, then details the output, then the read-only behavior, then filtering. It is longer than ideal but every sentence adds operational value. The bilingual opening is slightly redundant but not harmful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only reporting tool with no output schema, the description covers the key output elements, the data source, the read-only guarantee, and the main parameters. It does not explain the format parameter or the exact meaning of as_of, but the core call semantics are sufficiently complete for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain the purpose of include_detail and the filter dimensions (status, supplier, match date), which maps to status, supplier_ids, and from_date/to_date. However, it does not explain as_of, format, or workspaceId semantics, and the schema itself provides no descriptions. The description partially compensates but leaves several parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reports the status and exceptions of the three-way match (Belegabgleich), including summary counts, matched value by status, and per-line exception details. It names the specific resource (I04 three_way_match active rows) and distinguishes itself from sibling tools like match_three_way_exceptions and match_three_way_get by describing its aggregate/exception-reporting behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what the tool reads (live I04 three_way_match active rows), what it does not do (never auto-fixes a match), and how to filter (by status, supplier, match date) and expand detail (include_detail). It does not explicitly name alternative tools for when to use them instead, but the read-only, reporting nature is clear enough for an agent to select it over mutation tools like match_three_way_override or match_three_way_reverse.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
procurement_open_commitmentsARead-only
Offene Bestellverpflichtungen (open purchase commitments): every still-open PO line with its ordered, received and billed quantities and the residual commitment value (ordered minus billed, at the PO base price in integer CHF Rappen), plus aging-bucketed totals (0-30 / 31-60 / 61-90 / 90+ days). Cancelled POs contribute nothing and closed POs are omitted unless includeClosed. Pass group_by (supplier / item / aging) for a rollup, or leave it flat with cursor pagination (default 500, max 5000 rows). A DERIVED read: it computes over the live D02 po_line counters and posts nothing. Empty answers ok with zero totals. A supplier or item id from another workspace simply matches nothing (tenant isolation).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No | ||
| filter | No | ||
| format | No | ||
| group_by | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description explicitly states it is a DERIVED read that posts nothing, how cancelled and closed POs are treated (unless includeClosed), that empty answers return zero totals, and that cross-workspace ids match nothing due to tenant isolation. This is rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause earns its place: scope, quantities, aging buckets, aggregation options, pagination, derived-read safety, empty results, and tenant isolation. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without an output schema, the description still conveys the key output (line-level quantities, residual value, aging totals) and critical edge cases. It is nearly complete, though explicit filter/format parameter details would make it fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates substantially: it explains group_by values (supplier/item/aging), cursor pagination (default 500, max 5000), includeClosed behavior, and tenant-isolation semantics for filter ids. It does not fully describe the filter object shape or format parameter, leaving a small gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific product: every still-open PO line with ordered, received, billed quantities and residual commitment value, plus aging-bucketed totals. It distinguishes itself from report siblings by specifying derive-over-live-counters semantics, optional group_by rollups, and the flat cursor-paginated mode.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context is provided: use it for open purchase commitments, pass group_by for a rollup, or leave it flat for paginated line data, with defaults and limits stated. It does not explicitly name alternatives or say when not to use it, but the usage context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
procurement_po_cycleARead-only
PO-Durchlaufzeiten (PO cycle-time metrics) over POs that reached a terminal state (received / closed) in the window: the average and the p50 / p90 days for order-to-first-receipt and order-to-full-match, overall or grouped by supplier. Only completed POs count; open POs are excluded from the averages. Measures the timestamps the engine records (there is no sent_at stage, per the I06 reconciliation). A DERIVED read over the live PO / receipt / match documents; posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| to_date | No | ||
| group_by | No | ||
| from_date | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide readOnlyHint=true, and the description explicitly adds 'A DERIVED read over the live PO / receipt / match documents; posts nothing.' This confirms read-only behavior and adds that it computes over derived metrics, excludes open POs, and uses engine-recorded timestamps. For a read-only analytics tool with no destructive behavior, this is strong transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one fairly dense paragraph, packed with useful details: metric definitions, window, grouping, exclusions, timestamp source, and read-only nature. It is informative but a bit dense, and could have been structured with bullets or separate sentences for the exclusions and grouping. Still, every sentence adds meaning, so it earns a 4.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only metrics tool with 4 mostly self-explanatory parameters (workspaceId, from_date, to_date, group_by), the description covers the core semantics: what is measured, over what population, in what time frame, and why it is a derived read. It doesn't detail the output shape, but no output schema exists, and a mention of what the response contains (e.g., list of metrics with average/p50/p90) would improve it. Still, the description is quite complete for an analytics tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden for parameter semantics. It explains that from_date/to_date define the window, group_by presumably supports supplier grouping ('overall or grouped by supplier'), and the scope uses terminal states. It doesn't explicitly spell out each parameter name's format or allowed values, but it conveys the meaning of the date range and grouping role. Required workspaceId is mentioned in the description, so a 4 is warranted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool computes PO cycle-time metrics (average, p50, p90 days for order-to-first-receipt and order-to-full-match) over POs in a terminal state in a window, with grouping by supplier. This is a specific verb (compute aggregate metrics) + resource (POs in terminal state), and distinguishes it from siblings like po_list, po_get, po_open_lines, and other procurement_* analytics tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states when to use: when you need cycle-time metrics over completed POs in a date range, optionally grouped by supplier. It also explicitly excludes open POs, and mentions the absence of sent_at stage. It doesn't explicitly name sibling alternatives or say when not to use it, but the scope is clear enough to differentiate from transactional PO tools and other procurement reports.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
procurement_po_historyARead-only
Vollständige Bestellhistorie (complete PO history) for one po_id: an ordered timeline of creation, I01 revisions/amendments, posted I02 goods receipts, I03 landed-cost allocations and I04 three-way matches, each with its timestamp, actor, summary, amount / quantity and a deep-link ref. The running open quantity and open value are the live commitment (ordered minus billed) for the PO and reconcile to it exactly. A DERIVED read; posts nothing. A po_id from another workspace is not_found (tenant isolation); a just-created PO returns only its creation event.
| Name | Required | Description | Default |
|---|---|---|---|
| po_id | Yes | ||
| to_date | No | ||
| from_date | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial behavioral context beyond the readOnlyHint annotation: it explicitly states 'A DERIVED read; posts nothing', explains tenant isolation (not_found for other workspace), describes the running open quantity/value reconciliation, and notes that a just-created PO returns only its creation event. These details give the agent a clear model of the tool's side effects and edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a compact two-sentence block that is dense with specific information: event types, fields, running totals, derivation, and edge cases. No word is wasted, and the key purpose is front-loaded with 'complete PO history'. It earns its length by covering all essential aspects without repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only, single-PO history tool with no output schema, the description is remarkably complete: it lists event types, returned fields (timestamp, actor, summary, amount, quantity, deep-link), running totals, derivation, tenant isolation, and initial-state behavior. The only missing piece is the purpose of the optional to_date/from_date parameters, which keeps it from a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains po_id (one PO) and workspaceId (tenant isolation), but makes no mention of to_date or from_date, which are present in the schema and presumably filter the timeline. Two of four parameters remain undocumented both in schema and description, which is a significant gap for an agent trying to invoke the tool correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly specifies the tool returns the complete PO history for a single po_id, enumerating the exact event types (creation, I01-I04) and fields included. It is unambiguous about the resource and scope, and distinguishes itself from siblings by emphasizing the per-PO timeline rather than a list or current state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool: when you need the full ordered history of one PO (including revisions, receipts, allocations, and three-way matches). It does not explicitly name alternatives or exclusions, but the 'DERIVED read' and 'for one po_id' phrasing imply it is for read-only single-PO analysis, and the tenant isolation note clarifies cross-workspace behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
procurement_requisition_pipelineARead-only
Bestellanforderungs-Pipeline (requisition pipeline): the summary count and estimated value by status (draft / pending_approval / approved / partially_converted / converted / rejected / cancelled / closed), the open requisition rows with days-in-status, the linked PO and the conversion lag, and the conversion rate plus the average request-to-PO cycle over the window. Requisitions stalled beyond the workspace threshold are flagged. A DERIVED read over the live I00 requisition + conversion tables; posts nothing. Filter by status, requester or minimum aging.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| to_date | No | ||
| from_date | No | ||
| workspaceId | Yes | ||
| requester_ids | No | ||
| aging_days_min | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description states that this is a DERIVED read over the live I00 requisition and conversion tables and explicitly says 'posts nothing.' It also discloses the stalled-requisition flagging behavior, adding meaningful context about what the tool does with data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient: it packs numerous distinct outputs into a structured flow from summary metrics to open-line details to conversion analytics, and closes with read-only behavior and filters. The first sentence is long but every clause contributes new information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex analytics tool with no output schema, the description covers the key outputs, the underlying tables, side effects, and filter options. It doesn't specify the exact response structure or the definition of the workspace threshold, but the essential behavior is clear enough for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of explaining parameters. It maps the main filters (status, requester, minimum aging) and references the date window, but it does not explicitly explain workspaceId or the from_date/to_date parameters or their formats. Partial compensation, not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the exact resource (requisition pipeline) and enumerates the specific outputs: summary counts and estimated value by status, open lines with days-in-status, linked PO and conversion lag, conversion rate, and average cycle time. This level of detail makes it easy to distinguish from transactional siblings like requisition_get or procurement_po_cycle even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes it clear this is a read-only analytical view for the requisition pipeline, with filters by status, requester, or minimum aging, and explicitly says it posts nothing. This gives an agent a solid sense of when to use the tool, though it never names an alternative or states a specific when-not-to-use condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
procurement_spend_summaryARead-only
Einkaufsauswertung (procurement spend summary): grouped spend over the PO lines whose order date falls in the window, by supplier, item, item category or order month. Each row carries the document count and the ordered, received and billed value (the po_line counters times the base price, integer CHF Rappen); an optional compare_prior_period adds the prior equal window's billed value and the delta. All figures are DERIVED from the live PO lines and reproducible; nothing is posted. Empty range answers ok with zero totals. (Buyer and cost-centre groupings are not available: a PO carries neither, per the I06 spec reconciliation).
| Name | Required | Description | Default |
|---|---|---|---|
| filter | No | ||
| format | No | ||
| to_date | No | ||
| group_by | No | ||
| from_date | No | ||
| workspaceId | Yes | ||
| compare_prior_period | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses that figures are 'DERIVED from the live PO lines and reproducible', that 'nothing is posted', that an empty range returns zero totals, and that certain groupings are unavailable. This significantly enriches the agent's understanding of side effects, data provenance, and boundary behavior, far exceeding what the annotation alone conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but not bloated; every sentence contributes a distinct fact (grouping, metrics, optional compare, derived/non-posting, empty-range behavior, limitations). The opening repeats the tool name in German, which adds slight redundancy, but the information is front-loaded and logically ordered. Overall, it earns a high score for economy, though not a perfect 5 because of the translated-name duplication.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description compensates by describing row content (document count and ordered/received/billed values) and the optional compare field. It also explains data derivation, reproducibility, empty-range behavior, and grouping limitations. Still, the filter and format parameters are not described, and there is no mention of pagination or default grouping, leaving an agent to guess on those aspects in a tool with 7 parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry the parameter documentation. It does clarify the semantic roles of from_date/to_date (the 'window'), group_by (listing the allowed grouping values inline), and compare_prior_period (adds prior equal-window billed value and delta). However, it leaves filter and format entirely unexplainedaine. While it covers the most important parameters, the lack of any description for filter/format is a noticeable gap given the schema itself is silent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb–resource pair ('grouped spend over the PO lines') and specifies the grouping dimensions (supplier, item, item category, order month) and the metrics (document count; ordered, received, billed value). It is instantly distinguishable from procurement siblings like po_history or grir_clearing because it names the exact data source and aggregation scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is the tool for spend aggregation and explicitly states a limitation ('Buyer and cost-centre groupings are not available'), which serves as a 'when not to use' hint. However, it does not mention any alternative tools or provide explicit conditional routing among the many procurement-related siblings, so the guidance is left to inference rather than being spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
procurement_supplier_scorecardARead-only
Lieferanten-Scorecards (supplier scorecards) for a window: one row per supplier with the I05 overall score, on-time delivery %, average days late, price and quantity variance %, match-override and rejection rates, plus the live open commitment and in-period spend in Rappen. The metric math is DELEGATED to I05 (supplier_scorecard_get); this verb only adds the open commitment and spend and never writes or recalculates a score. Suppliers are the requested set, or every supplier with a PO in the window; those below min_activity are dropped and thin-data suppliers carry an insufficient_data warning. Posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| to_date | No | ||
| from_date | No | ||
| workspaceId | Yes | ||
| min_activity | No | ||
| supplier_ids | No | ||
| include_trend | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true, but the description goes further: it explicitly says the tool 'never writes or recalculates a score' and 'Posts nothing.' It also discloses supplier selection behavior (requested set vs. all suppliers with a PO), min_activity dropping, and insufficient_data warnings for thin-data suppliers. There is no contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: output shape, delegation boundary, supplier selection/filtering behavior, and a final side-effect guarantee. The key scoping and safety details are front-loaded before the implementation notes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description helpfully enumerates the returned columns and warns about insufficient_data. It also covers filtering and side-effect safety. Remaining gaps are minor: include_trend semantics, exact date formatting, and the threshold for min_activity are not specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden, and it compensates substantially: it explains the date window, supplier_ids selection, min_activity filtering, and the 'in-period spend' output semantics. However, include_trend is not described at all, and date parameter formats are only implied by 'window.'
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific resource and scope: one row per supplier with a precise list of metrics (I05 overall score, on-time delivery %, average days late, variance %, match-override/rejection rates, open commitment, in-period spend). It also distinguishes itself from the sibling supplier_scorecard_get by clarifying that metric math is delegated there and this verb only adds commitment/spend.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when this tool is appropriate: it adds open commitment and spend on top of I05 metrics and never writes or recalculates scores. It names the delegated supplier_scorecard_get, but it does not explicitly enumerate alternatives like supplier_performance_rank/trend or describe a when-not-to-use rule beyond the delegation relationship.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_budget_actualARead-only
Budget vs. Ist for one project, per project and per phase: budget, actual cost and hours, remaining, and the over-budget flag. COST ONLY, a pure read (P5) over the registered cost sources (B01 time, A17 bills, D02 purchases as they land; all zero until then): no revenue, no margin (that is B03), and nothing is ever written. includeSubprojects adds the base-currency rollup over the parent tree.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| projectId | Yes | ||
| workspaceId | Yes | ||
| includeSubprojects | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes far beyond the readOnlyHint annotation by disclosing that nothing is ever written, that cost sources are B01/A17/D02 and are zero until they land, and that includeSubprojects adds a base-currency rollup. This is rich behavioral context that the annotation alone does not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence adds distinct value: scope, metrics, exclusion, read guarantee, cost source behavior, and subproject behavior. It is front-loaded with the core purpose and avoids filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only reporting tool with no output schema, the description covers purpose, scope, metrics, data sources, zero-until-landed behavior, and the includeSubprojects flag. The main gap is the lack of any explanation of the asOf parameter, which could affect how current the returned data is.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries most parameter-semantics weight. It usefully explains includeSubprojects as a parent-tree rollup and makes projectId implied by 'one project', but it does not clarify the meaning or format of asOf, which is a significant omission.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the tool returns budget vs. actual cost and hours for a single project, broken out by phase, with remaining and an over-budget flag. The 'COST ONLY' and 'no revenue, no margin' statements distinguish it from related costing/reporting tools without requiring the agent to inspect schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for when it applies: cost-only budget/actual comparison over registered cost sources, explicitly excluding revenue and margin, which are attributed to B03. It does not name sibling tools explicitly, but the 'pure read' and scope exclusions give an agent enough to select it over alternative reporting tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_createB
Lege ein Projekt an (create a project master): name, client contact (C00), an editable auto-suggested code (P-0001 upward, unique per workspace), an integer-Rappen budget with budget hours, dates and an optional parent for a sub-project tree. A non-base currency snapshots the base-Rappen budget at the H-FX rate of the creation day (pass fxRate to assert one). Starts in status draft.
| Name | Required | Description | Default |
|---|---|---|---|
| code | No | ||
| name | Yes | ||
| endsOn | No | ||
| fxRate | No | ||
| currency | No | ||
| parentId | No | ||
| startsOn | No | ||
| contactId | Yes | ||
| budgetHours | No | ||
| budgetMinor | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the transparency burden. It does disclose useful behaviors: auto-suggested unique code, budget in integer Rappen, currency snapshot at creation-day rate, and initial draft status. However, it omits mutation safety details such as permission requirements, reversibility, or whether a successful create returns a project ID, so it is solid but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
All content earns its place and no filler exists, but the description is a single dense, run-on sentence mixing German and English. It is front-loaded with the action and then packs in many conditional behaviors, which makes it less scannable than well-structured shorter sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 12-parameter create tool with no annotations and no output schema, the description covers a lot of business logic but is not fully complete. It omits required-parameter explanations (workspaceId), idempotency semantics, return value, and any error conditions, so an agent still has uncertainty about a full successful invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description has to compensate. It does explain several parameters: name, client contact (C00), code format/uniqueness, budgetMinor, budgetHours, dates, parentId, currency, and fxRate. But it never mentions workspaceId or idempotencyKey, and it only vaguely refers to 'dates' instead of mapping to startsOn/endsOn, leaving gaps for 5 of 12 parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Lege ein Projekt an (create a project master)', a specific create-verb plus resource. It adds distinguishing details like optional parent for a sub-project tree and initial draft status, so an agent can tell it apart from project_update, project_set_status, or project_delete without guessing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what the tool does but never states when to use it versus alternatives, nor does it mention any exclusion criteria or give explicit 'when not to use' guidance. There is no referral to project_update for modifications or project_phase_add for child phases, leaving selection entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_deleteA
Hard-delete a DRAFT project that nothing references (no sub-project, no custom-field value, no linked file); its own draft phases go with it. Anything past draft has history and refuses with not_draft: real projects are closed, never erased (OR 957a spirit).
| Name | Required | Description | Default |
|---|---|---|---|
| projectId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: it reveals the destructive hard-delete nature, cascading deletion of draft phases, preconditions that must be checked, and the refusal error not_draft for non-draft projects. It also communicates the policy that real projects are closed, not erased, which is important behavioral context for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the primary action and immediately followed by the key constraints. Every clause adds useful behavioral or eligibility information, with no repetition of the tool name or generic filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a delete operation with no annotations and no output schema, the description thoroughly covers eligibility, side effects, and failure behavior. The only significant gap is the lack of explicit parameter-level guidance, especially around idempotencyKey, and the cryptic 'OR 957a spirit' reference adds little operational clarity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explain any of the three parameters: workspaceId, projectId, or idempotencyKey. The text mentions 'project' conceptually but gives no guidance on which IDs are required, how idempotencyKey should be used, or how the workspaceId relates to the project. The description fails to compensate for the schema's lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('hard-delete') and resource ('DRAFT project') while also naming the exact preconditions: no sub-project, no custom-field value, no linked file. It clearly distinguishes this from sibling project tools by limiting deletion to drafts only and explaining that draft phases are removed with the project.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use the tool: for DRAFT projects with no referencing objects. It also gives a clear exclusion by stating that anything past draft refuses with not_draft and real projects are never erased. It does not mention specific alternative sibling tools, but the usage boundary is strong enough for an agent to choose correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_getARead-only
Read one project in full, its phases embedded in sort order. phaseDone filters the phases to done (true) or open (false); savedViewId applies a saved view over the phase list (G00): its stored filters are merged underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| phaseDone | No | ||
| projectId | Yes | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint already set to true, the description adds valuable behavioral detail beyond the annotation: phases are embedded in sort order, phaseDone filters them, and savedViewId's stored filters are merged underneath explicit filters. This meaningfully explains how the tool behaves without contradicting the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no fluff. It front-loads the core purpose in the first sentence and packs the optional filter behavior into the second, with every word earning its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read operation, the description is complete: it states what is returned, how phases are ordered, and how the two optional filters behave. Without an output schema, the phrase 'read one project in full' gives the agent a sufficient mental model. Minor missing details like pagination or error conditions are not critical for a get operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description compensates by explaining phaseDone and savedViewId in detail, including the merge behavior. The remaining parameters, workspaceId and projectId, are required and self-explanatory from their names, so the description fills the most important gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb 'Read' and a clearly defined resource 'one project', immediately distinguishing it from sibling tools like project_list or project_phase_add. It also specifies that phases are 'embedded in sort order', which clarifies the exact return shape.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use this tool: to read a single project with its phases and optional filtering. It does not explicitly name alternatives such as project_list, so it lacks an explicit exclusion, but the intent is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_listARead-only
List the projects with code, status, budget and parent, filtered by status, contact, parent or a text query over code and name. savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| status | No | ||
| parentId | No | ||
| contactId | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description agrees. It adds useful behavioral context beyond the annotation by explaining how savedViewId works: stored filters are merged underneath explicitly named filters.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two well-organized sentences with no filler. The core list behavior and return fields are front-loaded, and the saved-view nuance is a separate second sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool, the description covers returned fields, filters, and saved-view behavior. Pagination, sorting, and valid status values are not mentioned, but these are minor for a simple filtered list operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden. It explains query as a text search over code and name, maps status/contact/parent to the relevant parameters, and describes savedViewId semantics. workspaceId is left to the schema/name, which is acceptable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'List the projects' with concrete returned fields (code, status, budget, parent). It is clearly distinguished from siblings like project_get and project_set_status by its collection semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: use this tool to list projects, with optional filtering by status, contact, parent, or text query. It does not explicitly name alternatives or when-not-to-use cases, but the intended usage is evident.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_phase_addA
Add a phase to a project (name, sort, integer-Rappen budget, budget hours, optional milestoneOn date). Refused on a closed project. A phase-budget sum exceeding the project budget answers ok with a phase_budgets_exceed_project warning, never a block: phase budgets are the plan, the project budget is the envelope.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| sort | No | ||
| projectId | Yes | ||
| budgetHours | No | ||
| budgetMinor | No | ||
| milestoneOn | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and adds non-obvious details: refusal on closed projects, the ok-with-warning behavior for budget overages, and the conceptual model that phase budgets are the plan while the project budget is the envelope. It does not cover return values or idempotency behavior, but the added context is substantive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences: the core operation and fields come first, then a refusal rule, then a nuanced budget-overage behavior. Every sentence contributes operational value with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write operation with no annotations and no output schema, the description covers the key invocation rules and domain-specific budget semantics. Missing idempotencyKey guidance and explicit return-behavior expectations are minor but real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies budgetMinor as an integer-Rappen budget, distinguishes budgetHours, and marks milestoneOn as optional. It does not explain sort or idempotencyKey semantics, though the required workspaceId and projectId are clear by name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Add a phase to a project') and enumerates the writable fields (name, sort, budget, hours, milestoneOn). This makes it clearly distinguishable from sibling tools like project_phase_update and project_phase_done just from the description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides an explicit precondition ('Refused on a closed project') and describes a relevant edge case (budget overage triggers a warning, not a block). It does not explicitly name alternatives such as project_phase_update for modifying existing phases, so it stops short of full when-to-use routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_phase_doneA
Mark a phase milestone reached (Erledigt): stamps doneAt (default today). A state assertion, so an already-done phase answers ok with alreadyDone true and the recorded date untouched, never a second stamp.
| Name | Required | Description | Default |
|---|---|---|---|
| doneAt | No | ||
| phaseId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so well: it discloses idempotency, the alreadyDone response field, the untouched recorded date, and the 'never a second stamp' guarantee. This goes well beyond what the schema alone provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The action is front-loaded, the default behavior is stated immediately, and the idempotency nuance earns its place in the second sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 4-param mutating tool with no annotations and no output schema, the description covers the main call path, default, defaults, and idempotent behavior. It stops just short of complete because it does not address idempotencyKey or describe the first-time success response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
It adds real meaning to doneAt (optional, defaults to today) and frames the whole operation as idempotent. However, with 0% schema description coverage, it never explains what idempotencyKey is for or what workspaceId/phaseId must reference, leaving a notable gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and object ('Mark a phase milestone reached') and immediately states the core side effect ('stamps doneAt'). Framing it as 'a state assertion' distinguishes it from project_phase_update and related project status tools even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context on when to call: when the phase is reached, with doneAt defaulting to today, and repeated calls are safe. It doesn't explicitly list alternatives or when not to use it, so it misses the top bar for explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_phase_updateA
Patch a phase (name, sort, budget, budget hours, milestoneOn). Refused on a closed project; carries the same phase-budget warning as project_phase_add.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | Yes | ||
| phaseId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description is the only behavioral disclosure. It mentions refusal on closed projects and a budget warning analogous to project_phase_add, adding real behavioral context. However, it omits other notable traits like auth requirements, return behavior, or side effects of partial updates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two crisp sentences, front-loaded with action and fields, followed by constraints. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers the core purpose and two key restrictions but leaves idempotency behavior, required identifiers' roles, and the shape of the patch object (beyond listed fields) undocumented. It suffices as a baseline but is not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% coverage, but the description lists valid patch keys (name, sort, budget, budget hours, milestoneOn), which is essential because the patch object schema is opaque. It does not explain workspaceId, phaseId, or idempotencyKey, leaving their semantics to inference from names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States verb 'Patch' + resource 'phase' and enumerates updatable fields (name, sort, budget, budget hours, milestoneOn). References sibling project_phase_add for warning context, differentiating it slightly, but does not explicitly contrast with project_phase_done or project_phase_add beyond that.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a condition ('Refused on a closed project') and references the same warning as project_phase_add, but no explicit guidance on when to use this tool versus creating or completing phases. Usage context must be inferred from the verb and field list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_set_statusA
Move a project through its lifecycle: draft to active, active to on_hold and back (Pausieren), active or on_hold to closed (Abschliessen), and the audit-logged reopen closed to active (Wieder öffnen). An illegal pair refuses with invalid_transition; closing runs the registered close guards (B01 adds the open-time guard when it lands).
| Name | Required | Description | Default |
|---|---|---|---|
| status | Yes | ||
| projectId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It goes beyond a generic 'update status' by disclosing that illegal transitions are rejected with invalid_transition, that closing triggers registered close guards, and that reopening is audit-logged. It does not cover authorization or return-value behavior, so it is not a 5, but it is unusually transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two information-dense sentences, front-loaded with the core action and state machine, then error and side-effect behavior. Every clause earns its place; the parenthetical German labels add precision without bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides the state transition rules, illegal-transition error name, close-guard side effects, and audit-logged reopen behavior — strong coverage for a state-change tool with no output schema. It is slightly incomplete because it does not describe what the tool returns or how idempotencyKey behaves, but these are secondary for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning to the status parameter by listing the concrete lifecycle values (draft, active, on_hold, closed). However, it says nothing about idempotencyKey, workspaceId, or projectId beyond their schema names, leaving the optional idempotency parameter's purpose undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a clear, specific action: moving a project through its lifecycle, and enumerates the exact allowed transitions (draft→active, active↔on_hold, active/on_hold→closed, closed→active reopen). This makes it immediately distinguishable from sibling project tools like project_update, project_delete, and project_get.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for when to use the tool by specifying the full state machine and noting invalid transitions. It does not explicitly name alternatives or when-not-to-use conditions, but the lifecycle framing effectively scopes its usage away from general project mutation tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
project_updateA
Patch a project master (name, contact, code, dates, parent, budget). Editable while draft, active or on_hold; a closed project refuses with project_closed. The H-FX base snapshot is re-taken only when currency or budgetMinor change, never silently re-rated.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | Yes | ||
| projectId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses meaningful behavioral traits: state-dependent editability, the specific error code project_closed for closed projects, and a non-obvious side-effect rule ('H-FX base snapshot is re-taken only when currency or budgetMinor change, never silently re-rated'). This goes well beyond a bare mutation description, though it does not address success responses or validation beyond the state check.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler. The primary action is front-loaded, followed by the state constraint and the side-effect rule. Every sentence contributes new information, and the structure makes the most important information (what and when) appears first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essentials: target resource, editable fields, valid lifecycle states, the refusal error, and an unusual side-effect. Given the tool is a patch/update with no output schema, the description is reasonably complete for invocation. The main gap is a description of the return value, but for a patch tool this is a modest omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description must compensate. It lists the top-level patch keys (name, contact, contact, code, dates, parent, currency, budgetMinor), which is useful. However, it does not define the shape of nested values like 'dates' or clarify the patch object structure, and workspaceId/projectId/idempotencyKey are left entirely to inference. The description adds some field-level semantics but not enough to fully bridge the missing schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource, 'Patch a project master', followed by a parenthetical list of editable fields (name, contact, code, dates, parent, budget). This clearly differentiates the tool from siblings like project_set_status (status changes) and project_delete (removal), so an agent can immediately understand what this tool is for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when the tool applies: 'Editable while draft, active or on_hold; a closed project refuses with project_closed.' This provides both when and when-not usage context. It does not explicitly name alternatives (e.g., project_set_status for status changes), so it stops short of a 5, but the clear state-based guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
propose_dunning_runA
Propose a Mahnlauf: read the overdue open items from the OP-Liste (A16) as of a date, assign each invoice its next level (1./2./3. Mahnung; an invoice whose 3. Mahnung is issued is terminal), group per debtor, and persist a reviewable DRAFT run. Nothing is booked, rendered or sent: this is the safe half, and the reason an agent or an automation rule may call it without a confirmation. One run per day: re-proposing the same day returns the existing run. A fully paid invoice never appears; a partly paid one appears with its reduced open amount. asOf may look back but never forward. CONSEQUENCE: Freezes a reminder run proposal over the open items as of the given date.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and succeeds: it discloses side effects (persists a DRAFT run), safety (nothing booked/rendered/sent), filtering (fully paid excluded, partially paid with reduced amount), and the terminal nature of a 3rd-level reminder. It also states the consequence ('freezes a reminder run proposal') which adds clarity beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence adds meaningful information: definition, safety, idempotency, exclusions, temporal constraint, and consequence. It is front-loaded with the main action and avoids filler. The compact structure suits the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and a sparse input schema, the description is largely complete for operational understanding: it covers purpose, constraints, side effects, and safety. It does not mention how to retrieve the created draft run (e.g., via list_dunning_runs) or what the returned object looks like, but that is a minor gap when siblings exist and no output schema is defined.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains 'asOf' semantically ('may look back but never forward') and alludes to idempotency behavior (re-proposing same day returns existing run), but it does not explicitly connect the 'idempotencyKey' parameter to that behavior, and 'workspaceId' is entirely unexplained. This leaves the meaning of two of three parameters partially inferred rather than clearly stated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Propose') with a clear resource ('Mahnlauf'), source ('OP-Liste (A16)'), and processing steps (assign levels, group per debtor, persist DRAFT). It explicitly differentiates itself from related tools by stating 'Nothing is booked, rendered or sent' – distinguishing it from issue_dunning_run and send_dunning_run in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly indicates this is the safe, non-finalizing step and may be called without confirmation, implying when to use it. It also states the idempotency rule ('re-proposing the same day returns the existing run') and the asOf constraint ('may look back but never forward'). However, it does not explicitly name an alternative tool for the actual booking/sending step, leaving the agent to infer from sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
provision_createA
Describe a Rückstellung (OR Art. 960e) and save it as a DRAFT: reason is one of the eight statutory ids (Garantie, Ferien und Überzeit, Prozess, Grossreparatur, Sanierung, Restrukturierung, Steuern, Sonstige; invalid_reason names the exact ids, and sonstige needs a description of at least 10 characters); provisionAccount is 2330, 2600 or a liability numbered 23xx/26xx; expenseAccount is the P&L account the formation is charged to. Returns the draft with its lines (Dr expense / Cr provision). Nothing is posted.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| periodEnd | Yes | ||
| amountMinor | Yes | ||
| description | Yes | ||
| workspaceId | Yes | ||
| expenseAccount | Yes | ||
| idempotencyKey | Yes | ||
| provisionAccount | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden, and it does so thoroughly. It discloses the DRAFT-only side effect, the return shape (Dr expense / Cr provision lines), and validation rules for reasons and accounts. This goes well beyond what the schema alone would communicate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the core behavior, and every clause adds domain value. It is slightly run-on with nested parentheticals, but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and eight required parameters, this description is unusually complete: it explains the domain concept, legal reference, account constraints, side effects, and return shape. Minor gaps remain around the format or units of amountMinor, periodEnd, and the exact semantics of idempotencyKey, but overall the agent has enough context to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description adds crucial parameter meaning beyond the bare schema: exact statutory reason ids, account number constraints, expense account semantics, and the description length requirement. It leaves periodEnd, amountMinor, workspaceId, and idempotencyKey unspecified, but the most restrictive and domain-specific parameters are explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it describes a Rückstellung provision and saves it as a DRAFT, explicitly clarifying the operation's scope. The final sentence 'Nothing is posted' clearly distinguishes it from posting-related provision siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description conveys clear context: this tool creates a draft rather than posting it, and explains the draft's resulting lines. It does not name alternative tools like provision_post or provision_post, but 'Nothing is posted' gives the key routing signal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
provision_discardA
Retire a drafted Rückstellung before it posts. The row stays on record as discarded with its reason; a posted provision cannot be discarded (already_posted), reverse or release it instead.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| provisionId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. It explains the mutation effect (row stays on record as `discarded`), that the reason is retained, and the error condition (`already_posted`). It could additionally mention irreversibility or idempotency behavior, but it covers the most important state change and constraint well.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loads the core action and condition, and includes technical details (state value, error code, alternatives) without waste. Every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple lifecycle action with no output schema and no annotations, this description is largely complete: it defines scope, outcome, error handling, and alternatives. It omits exact reverse/release tool names and idempotency semantics, but for this level of complexity the agent has enough to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only adds meaning for `reason` by indicating that the row retains the reason; `provisionId`, `workspaceId`, and especially `idempotencyKey` are left unexplained. The parameter names are somewhat self-evident, but the description does not clarify formats, requirements, or the purpose of the idempotency key.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Retire') and a specific resource ('a drafted Rückstellung') with a clear temporal condition ('before it posts'). It also distinguishes itself from related operations by stating that posted provisions must instead be reversed or released, so the tool's role is unambiguous even among many similar siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when this tool is appropriate ('drafted ... before it posts') and when it is not ('a posted provision cannot be discarded'). It directly names the alternative behavior—'reverse or release it instead'—mapping to sibling tools like provision_reverse and provision_release, giving the agent clear routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
provision_getARead-only
Read one Rückstellung: the row, its formation lines, every release with its entry (and the reversal of that entry, when one exists) and the open balance, derived as the formation amount minus the live releases.
| Name | Required | Description | Default |
|---|---|---|---|
| provisionId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already declares the operation is read-only. The description adds value by explaining the response composition: the row, formation lines, releases with entries, reversals, and the derived open balance. It also discloses the calculation of the open balance (formation amount minus live releases), which is beyond the annotation. No contradiction with the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that starts with the core purpose ('Read one Rückstellung') and then enumerates the returned components in a logical order. It is concise with no filler, and each phrase adds meaningful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read operation with two string parameters and no output schema, the description covers the essential return contents and even explains how the open balance is derived. It does not address edge cases like not-found errors or pagination, but those are minor for a get-by-id tool. The description is sufficiently complete for an agent to know what to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has no descriptions for workspaceId or provisionId, and the description does not mention either parameter. It neither explains that provisionId identifies the specific provision nor that workspaceId scopes the request. With 0% schema coverage, the description fails to compensate, leaving the agent without any parameter semantics beyond the names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the specific verb 'Read' and the resource 'one Rückstellung', and enumerates exactly what is returned: the row, formation lines, releases with entries, reversals, and open balance. It clearly differentiates from sibling tools like provision_list (which lists provisions) and mutation tools (provision_create, provision_post, etc.) by focusing on a single read with detailed content.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is the detailed single-record read for provisions, as opposed to provision_list for listing. It provides clear context that this returns all associated data for one provision, but it does not explicitly name alternatives or state 'use this instead of X'. Since the purpose is unambiguous given the sibling set, a score of 4 is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
provision_listARead-only
List Rückstellungen, newest period first, each with its open balance, filtered by periodEnd, status (draft | posted | released | reversed | discarded) or reason. openTotalMinor sums the open balances. savedViewId applies a saved view (G00).
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| status | No | ||
| periodEnd | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the default ordering ('newest period first'), the computed `openTotalMinor` aggregation, and the effect of `savedViewId` — beyond the readOnlyHint annotation. It doesn't mention pagination or response details, but for a read-only list tool the annotation plus this context is sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences front-load the core purpose, then add only the essential filter/behavior details. Every clause serves a purpose; no filler or repetition of schema names.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list with five string params and no output schema, the description covers purpose, ordering, filters, status values, per-item balance, and total aggregation. Missing pagination details are a minor gap given the simple list semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden. It explains periodEnd, status (with explicit allowed values), reason, and savedViewId (G00 behavior); only workspaceId, a self-evident required scope parameter, is not elaborated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource (provisions/Rückstellungen), the operation (list), the sort order (newest period first), and the data shape (open balance). This clearly separates it from sibling provision_create/post/release/reverse/get tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the filterable dimensions (periodEnd, status, reason) and that savedViewId applies a saved view (G00), implying when to call it: to retrieve a filtered list of provisions. It does not explicitly name alternatives or exclusions, but the list-vs-get contrast is evident from the sibling set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
provision_postA
Post a drafted Rückstellung as ONE entry (Dr expense / Cr provision, source provision) dated periodEnd. No automatic reversal: OR 960e Abs. 4 does not release a provision by the calendar. Idempotent per provision; a different key on a posted one is already_posted. CONSEQUENCE: Posts the provision as an immutable entry; it is released or reversed later, never edited.
| Name | Required | Description | Default |
|---|---|---|---|
| provisionId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the full burden. It discloses idempotency, the `already_posted` conflict code, immutability ('posted as an immutable entry; it is released or reversed later, never edited'), and the lack of automatic reversal based on time ('OR 960e Abs. 4 does not release a provision by the calendar'). This is excellent behavioral transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core action, then key behavioral caveats, then a consequences summary. It is dense but not bloated. The 'CONSEQUENCE:' label is a slightly formal flourish, but the sentences all earn their place. Minor redundancy between 'no automatic reversal' and the consequence line is acceptable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a financial posting action with no annotations and no output schema, the description covers the accounting effect, source, date, idempotency, immutability, and the error code for double posting. It doesn't describe the success response or the returned structures, but since there is no output schema, the agent may still lack that info. However, given the complexity of the tool and the strong behavioral caveats, it is quite complete, with minor gaps around response/return and explicit permission or precondition checks.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for the three parameters. It gives some meaning to `idempotencyKey` via idempotency semantics and to `provisionId` by calling it 'a drafted Rückstellung', but it doesn't explicitly explain `workspaceId` or the format/constraints of the idempotency key. The description adds value but leaves some parameter semantics to the schema or agent inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Post a drafted Rückstellung as ONE entry (Dr expense / Cr provision, source `provision`) dated `periodEnd`.' This clearly distinguishes the tool from the many related siblings (provision_release, provision_reverse, provision_discard, post_entry, etc.).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use it: to post a drafted provision, with a one-entry accounting booking and an idempotency key. It gives a conflict signal (`already_posted`) that tells the agent when a provision is already posted. It doesn't explicitly name alternatives or exclusions, but context makes the usage fairly clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
provision_releaseA
Release part or all of a posted Rückstellung: posts Dr provision / Cr targetAccount (the original expense account, or an income account) for amountMinor on date. Refuses release_exceeds_balance above the open balance; the provision reads released at zero. Idempotent on idempotencyKey for as long as the release that key booked still stands; a key whose release was undone refuses already_reversed_key (release again under a new key). A release is undone only with provision_release_reverse; the raw reverse_entry refuses its entry with owned_by. CONSEQUENCE: Posts the release of part or all of the provision against the target account; a release is undone only by a reversal.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | ||
| amountMinor | Yes | ||
| provisionId | Yes | ||
| workspaceId | Yes | ||
| targetAccount | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: it discloses a mutating posting (Dr provision / Cr targetAccount), idempotency semantics tied to idempotencyKey 'for as long as the release that key booked still stands', and three concrete refusal codes (release_exceeds_balance, already_reversed_key, owned_by). It also states the postcondition that the provision 'reads released at zero', exactly the behavioral context annotations would otherwise provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action, error states, idempotency, and reversal path are presented in a logical, front-loaded order with high information density. The trailing CONSEQUENCE paragraph largely restates the first sentence ('Posts the release... undone only by a reversal'), adding redundancy without new information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a complex financial mutation (6 required parameters, zero annotations, no output schema, 0% schema coverage), yet the description covers the accounting effect, all refusal conditions, idempotency, the undo path, and the sibling relationship — everything needed to select and invoke the tool correctly. The only omissions (response shape, explicit permission requirements) are minor and do not impede correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds real meaning to targetAccount ('the original expense account, or an income account') and to idempotencyKey (idempotency behavior, when a key is refused), and weaves date/amountMinor into the posting sentence. workspaceId and provisionId are only implied, though both are reasonably inferable from the operation context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Release part or all of a posted Rückstellung' with the exact accounting entry 'posts Dr provision / Cr targetAccount'. The Dr/Cr detail plus the error semantics distinguish this from siblings like provision_post, provision_reverse, and provision_release_reverse without opening schemas. The untranslated German 'Rückstellung' is minor friction, but the context makes the meaning clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the undo path: 'A release is undone only with provision_release_reverse; the raw reverse_entry refuses its entry with owned_by.' It also states when a key is refused (already_reversed_key) and the required retry rule ('release again under a new key'), telling the agent which tool and which key to use in each state. This is concrete when/when-not guidance naming sibling alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
provision_release_reverseA
Undo one release of a Rückstellung with a real reversing entry (dated the release date unless date names another open day), so the open balance and the posted / released reading follow from the journal again. This is the ONLY way to reverse a release entry: the raw reverse_entry refuses it with owned_by, and so it refuses the undo itself. Refuses not_found and already_reversed. Idempotent on idempotencyKey. CONSEQUENCE: Posts a reversing entry against one release of the provision, restoring its open balance; it cannot be un-posted.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | ||
| reason | No | ||
| releaseId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full responsibility and discloses key behavior: it posts a real reversing entry, restores the open balance, is idempotent on `idempotencyKey`, and 'cannot be un-posted'. It also states the effect on `posted`/`released` readings, making consequences transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but not overly long. It front-loads the main action and then provides necessary caveats (refusals, idempotency, consequence). Every sentence adds value, though the phrasing 'and so it refuses the undo itself' is slightly redundant with the preceding clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (financial reversal, idempotency, refusal conditions, non-reversibility), the description covers the main aspects. It mentions the effect on balance, posting behavior, and error cases, but does not explicitly mention permissions or prerequisites. However, these are implied by the refusal conditions and the domain context, so the completeness is adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has no parameter descriptions (0% coverage), so the description must compensate. It explains `date` ('dated the release date unless `date` names another open day') and `idempotencyKey` (idempotent), but does not explain `reason` or the obvious identifiers. Since `releaseId` and `workspaceId` are self-evident, the gap is minor, but `reason` could benefit from a note.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action: 'Undo one release of a Rückstellung with a real reversing entry.' It specifies the resource (a release of a provision) and the mechanism (a reversing entry), and distinguishes it from the raw `reverse_entry` by noting it is the ONLY way to reverse a release entry. This differentiates it clearly from siblings like `provision_reverse` and `reverse_entry`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'This is the ONLY way to reverse a release entry' and names the alternative (`reverse_entry`) that refuses with `owned_by`. It also lists refusal conditions (`not_found`, `already_reversed`) and idempotency on `idempotencyKey`, giving clear when-to-use and error-handling guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
provision_reverseA
Reverse the formation entry of a posted Rückstellung (a real reversing entry, dated periodEnd unless date names another open day). Refuses release_blocked while any release stands unreversed (undo the releases first through provision_release_reverse, newest first) and already_reversed a second time. This is the ONLY way to reverse the formation: the raw reverse_entry refuses it with owned_by. CONSEQUENCE: Posts a reversing entry against the provision formation; it cannot be un-posted.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | ||
| reason | No | ||
| provisionId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully discloses the behavioral profile: it is 'a real reversing entry', cannot be un-posted, refuses a second reversal, and rejects while releases stand. It even surfaces the error conditions `release_blocked`, `already_reversed`, and `owned_by`, which gives the agent accurate expectations before invoking a destructive operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient: every clause adds a distinct fact—operation and date, refusal conditions with remedy, exclusivity, and irreversible consequence. The core behavior is front-loaded, and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description supplies the essential operational context: prerequisite cleanup steps, the only applicable alternative, error conditions, default posting date, and the fact that the result cannot be reversed. An agent has enough information to select the tool correctly and anticipate the effect of calling it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, and the description adds meaningful semantics mainly for `date` ('dated `periodEnd` unless `date` names another open day') and ties `provisionId` to the formation being reversed. However, it does not explain `idempotencyKey`, `reason`, or `workspaceId` beyond their self-evident names, so the coverage gap is only partially closed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and object: 'Reverse the formation entry of a posted Rückstellung'. It also distinguishes itself from `reverse_entry` and `provision_release_reverse`, stating 'This is the ONLY way to reverse the formation', so an agent can tell exactly what resource this tool acts on.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when not to call it: it refuses with `release_blocked` while releases are unreversed and instructs the agent to undo them via `provision_release_reverse` first, newest first. It also names the alternative raw `reverse_entry` and explains why it cannot be used, giving clear routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quotes_acceptA
Nimm eine Offerte an: either by the single-use token (the cloud accept host, the one write reachable without an authenticated actor, scoped to exactly one quote by its H-TENANT token-hash lookup) or by quoteId + actor (the human "Als angenommen markieren"). Refuses acceptance after valid_until (quote_expired, OR Art. 3), stamps accepted_by, invalidates the token (single-use), and moves the quote to accepted. Posts nothing; conversion is a separate step.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| token | No | ||
| quoteId | No | ||
| acceptedBy | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It covers expiration refusal, accepted_by stamping, single-use token invalidation, state transition to accepted, and the unauthenticated write path. This is unusually rich disclosure for a write tool, especially without annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one long, dense run-on paragraph mixing German and English with heavy parentheticals ('the cloud accept host, the one write reachable without an authenticated actor...'). It contains valuable information but is poorly structured and harder to parse than it should be for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a state-changing tool with no output schema and six parameters, the description covers the critical behavioral context: the two invocation modes, expiry handling, side effects, and the boundary with conversion. It is missing clarification of workspaceId and idempotencyKey, but the core call semantics are sufficiently complete to guide selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and several parameters are opaque, but the description adds real meaning to token, quoteId, actor, and acceptedBy by explaining the token-hash lookup, the human marking flow, and the accepted_by stamp. It does not explain workspaceId or idempotencyKey, which are important gaps given zero schema coverage, so it does not fully compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the operation explicitly ('Nimm eine Offerte an' / accept a quote) and distinguishes it from sibling quote actions by specifying the two acceptance paths and noting that conversion is a separate step. It is clearly scoped to acceptance only, which separates it from quotes_convert, quotes_decline, and quotes_revise.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete selection criteria: use the single-use token for the unauthenticated cloud accept flow, or quoteId + actor for the human-facing 'Als angenommen markieren' action. It also states the negative condition 'Posts nothing; conversion is a separate step', which helps an agent understand what this tool is not for. It does not name sibling tool names explicitly, but the boundary conditions are clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quotes_convertA
Wandle eine angenommene Offerte um in einen Auftrag (order) oder direkt in eine Rechnung (invoice): a DIRECT 1:1 pass-through to A10 convertDocument, which clones the frozen document_line rows (the VAT trace travels with them) and marks the source converted. C02 posts nothing; revenue is recognised only when A11 later issues the invoice. Idempotent and single-issue: converting an already-converted quote returns the existing target, never a second document. A non-accepted quote is refused (illegal_transition).
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | ||
| quoteId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does an excellent job: it discloses the effect on document_line rows, VAT trace propagation, source marking, absence of posting, revenue recognition timing, idempotency guarantees, and the illegal_transition error for non-accepted quotes. This is rich, actionable behavioral information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the primary purpose, followed by key behavioral notes. However, it includes internal module references (A10, C02, A11) that are unexplained and could add noise for an agent trying to parse meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the operation's complexity, it covers the critical preconditions, side effects, and invariants. It does not mention the response payload or how to specify the target 'to' value programmatically, but these are secondary given no output schema and the clear prose target description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, leaving the description to explain all parameters. It only implicitly explains 'to' (order/invoice) and 'quoteId' (the source), but offers no guidance on 'workspaceId' or 'idempotencyKey' despite emphasizing idempotency. An agent would be unclear how to supply the idempotency key or whether it is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (convert) and resource (accepted quote) with two target types (order/invoice). However, it does not distinguish from similar siblings like sales_order_from_quote or the underlying convert_document, leaving ambiguity about when this specific tool is preferred.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It specifies the precondition (accepted quote) and idempotency behavior, and explicitly refuses non-accepted quotes. But it does not advise when to use this vs. the sibling convert_document or sales_order_from_quote, and mentions a pass-through to A10 convertDocument without clarifying that this is the correct high-level entry point.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quotes_createB
Lege eine Offerte an: a priced offer on a C00 contact. Each line resolves its price once through D00 (contact list, segment list, item base) and its tax code once through the item A05 default, then snapshots both as literals on the shared document_line row (the freeze). Rides A10 createDocument (type quote), always a draft, and POSTS NOTHING (quote has a no-op poster). valid_until is the OR Art. 3 binding window and must be today or later; deal_id links a C01-seeded quote. Foreign contact/item ids are refused (invalid_reference, H-TENANT).
| Name | Required | Description | Default |
|---|---|---|---|
| intro | No | ||
| lines | No | ||
| outro | No | ||
| dealId | No | ||
| currency | No | ||
| contactId | No | ||
| validUntil | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers substantial context: the freeze semantics (resolves price/tax once and snapshots them), always-draft status, no-op poster, valid_until binding constraint, and rejection of foreign IDs with error codes. This goes well beyond a simple create statement, though it leaves idempotency and response behavior undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a dense, jargon-heavy block mixing German and internal codes (C00, D00, A05, A10, OR Art. 3). It front-loads the purpose but then meanders into implementation details and constraints without clear sections or glossaries. It could be split into purpose, behavior, and validation, and translated for a broader audience.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 9 parameters, no output schema, and no annotations, the description is not complete enough. It covers creation semantics and a few constraints, but omits what the API returns, how to structure the lines array, what each property means, and what internal codes refer to. An agent without domain knowledge would struggle to make a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 9 parameters and 0% description coverage, so the description must compensate. It explains only valid_until (must be today or later), deal_id (links a C01-seeded quote), and foreign ID rejection. Parameters like intro, outro, currency, lines, and idempotencyKey are left unexplained, making it insufficient for an agent constructing a correct call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Lege eine Offerte an' / create a quote) and clarifies it is a priced offer on a C00 contact, always a draft. It distinguishes from quote manipulation tools by noting the draft state and 'no-op poster,' but does not explicitly name sibling alternatives and relies on heavy internal jargon.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies when to use it: for creating a draft quote, not for updating or sending. Mentions 'Rides A10 createDocument' and 'always a draft,' which signposts the creation scope. But it lacks explicit when-not-to-use guidance and does not reference specific sibling tools like quotes_update or quotes_send.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quotes_declineA
Lehne eine Offerte ab (by quoteId or token): moves the quote to declined, persists an optional decline_reason, invalidates the token, and logs the timeline note (OP5). A declined quote can only be superseded by a revision.
| Name | Required | Description | Default |
|---|---|---|---|
| token | No | ||
| quoteId | No | ||
| workspaceId | Yes | ||
| declineReason | No | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses several behavioral traits: moves the quote to declined, persists an optional decline_reason, invalidates the token, and logs a timeline note (OP5). It also states the post-condition that a declined quote can only be superseded by a revision. This is good behavioral disclosure for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each carrying distinct information: the action and its effects, the optional reason, and the post-condition. No filler or repetition. The most important information (what the tool does) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers the key behavioral effects, the identifier options, and the post-condition. It doesn't describe the return value or error cases, but the core information an agent needs to invoke it correctly is present. The missing parameter semantics for workspaceId and idempotencyKey are a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the two alternative identifiers (quoteId or token) and mentions decline_reason as optional. However, it doesn't explain workspaceId, idempotencyKey, or the exact format/constraints of the parameters. The description adds some meaning but leaves several parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (decline), the resource (a quote), and the mechanism (by quoteId or token). It also distinguishes itself from the sibling quotes_revise by noting that a declined quote can only be superseded by a revision. This is a specific, unambiguous purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when declining an offer, and it mentions the alternative path (revision) for superseding a declined quote. It doesn't explicitly say 'use quotes_revise instead when you want to modify a quote', but the context is clear enough for an agent to select this tool over siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quotes_expire_sweepA
Verfallene Offerten kehren aus der offenen Liste: for every sent quote past its valid_until, transition it to expired so lists and forecasting never count a dead offer. Agent-scheduled; the Studio list triggers the same verb lazily. Returns { expired: N }. A no-op when nothing is past validity.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that the tool mutates quote states (transitions to expired), returns a count { expired: N }, and is a no-op when nothing is past validity. It also explains the business rationale (preventing dead offers from being counted). It does not mention idempotency, permissions, or error conditions, but the core behavior is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description begins with a German phrase that adds little value and could confuse non-German-speaking agents. The English portion is informative but not tightly structured; the key purpose is stated, but the German preamble and the explanatory clause about lists and forecasting could be trimmed for clarity. It is not as concise as it could be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there is no output schema, the description explicitly states the return format ({ expired: N }) and the no-op case. It specifies the exact condition for expiration (sent quotes past valid_until) and the downstream effect. Missing details include error handling, rate limits, or whether the operation is scoped to a single workspace, but the core execution context is adequately described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description does not mention workspaceId or idempotencyKey at all. The agent must infer that workspaceId scopes the operation and idempotencyKey ensures safe retries. This is a significant gap since the description should compensate for the lack of schema details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: transition sent quotes past their valid_until to expired. It specifies the resource (quotes), the condition (past valid_until), and the outcome (expired status). This distinguishes it from siblings like quotes_revise or quotes_update, which handle individual quote modifications rather than bulk expiration sweeps.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description notes that it is 'Agent-scheduled' and that the Studio list triggers the same verb lazily, implying it is intended for scheduled, automated use rather than manual invocation. However, it does not explicitly state when to avoid using it or mention alternative tools for individual quote expiration, leaving the routing decision partially to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quotes_getARead-only
Lies eine Offerte mit Positionen, Statusverlauf und den C02-Feldern (validUntil, version, dealId, intro/outro, acceptedBy, declineReason). A sent quote past its validity is flagged expired on read (derived, never mutating; the sweep does the durable transition).
| Name | Required | Description | Default |
|---|---|---|---|
| quoteId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint=true annotation, the description discloses a subtle behavioral trait: an expired sent quote is flagged expired on read, explicitly described as derived and never mutating, with the durable transition delegated to the sweep. This is precisely the non-obvious context an agent needs to avoid assuming the read itself changes state, and it reinforces rather than contradicts the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler: the first front-loads the action and content, the second adds a single behavioral nuance that earns its place. The field list is compact and immediately useful, and no information is repeated from annotations or the input schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value explanation, and it does cover what is returned: positions, status history, and the enumerated CO2 fields, plus the derived expiry-flag behavior. It is missing only minor conveniences such as not-found/error behavior or explicit parameter guidance, which keeps it from a perfect score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of explaining the parameters, but it explains neither workspaceId nor quoteId. The field names are self-evident and the German 'eine Offerte' loosely implies quoteId identifies the quote, yet nothing clarifies the scoping role of workspaceId or any format constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource — 'Lies eine Offerte' (read a quote) — and enumerates the returned content: line items, status history, and the named CO2 fields. The read semantics and 'never mutating' statement clearly separate it from mutating siblings like quotes_send, quotes_accept, quotes_decline, and quotes_revise, and the single-object scope distinguishes it from quotes_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: an agent can infer this tool is for retrieving full quote details, but the description never explicitly says 'use this when you need a quote's items and status' or names alternatives like quotes_list or quotes_get vs. quotes_revise. No when-not-to-use or exclusion conditions are given, so the agent must derive the context from the verb 'Lies' and the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quotes_listARead-only
Liste die Offerten (P5): filter by status or contact. Only the newest version of a chain shows by default (a superseded quote is hidden unless includeSuperseded); each row carries validUntil, version and a derived expired flag. savedViewId applies a saved view (G00): its stored filters merge underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| contactId | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| includeSuperseded | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses important behaviors: only the newest version shows by default, superseded quotes are hidden unless includeSuperseded is set, each row carries validUntil, version, and an expired flag, and savedViewId merges stored filters under explicit ones. This is rich behavioral context that aids correct invocation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the primary purpose, then behavioral details. No fluff; every clause adds value. The structure is efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a list tool with five parameters and no output schema, the description covers the key operational aspects: filtering, versioning, and saved-view merging. It lacks explicit mention of pagination or output structure, but these are less critical for a list operation. Overall, it provides enough to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameters. It covers status, contactId, includeSuperseded, and savedViewId, but omits workspaceId, which is required. While workspaceId is self-explanatory, the description does not explicitly mention it. It also doesn't detail return format, but that's beyond parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists quotes (Offerten) and specifies filtering by status or contact. It also distinguishes this list operation from sibling tools like quotes_get by describing the versioning and saved-view behavior unique to listing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides context on how to use the tool (filtering by status or contact, including superseded quotes, applying saved views) but does not explicitly mention when not to use it or point to alternatives like quotes_get for single-quote retrieval. The usage guidance is implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quotes_reviseA
Revidiere eine gesendete/abgelehnte/abgelaufene Offerte: clones the FROZEN lines byte-for-byte into a new draft (version n+1, supersedes_id at the old row) and retires the old version as superseded, invalidating its token. A draft is edited in place, not versioned (edit_draft_instead); an accepted or converted quote is the binding record and is refused (illegal_transition).
| Name | Required | Description | Default |
|---|---|---|---|
| quoteId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With zero annotations, the description carries the full burden of behavioral disclosure, and it delivers richly: byte-for-byte cloning of frozen lines, version n+1 with supersedes_id on the old row, retirement of the old version as superseded, token invalidation, in-place draft editing, and rejection of accepted/converted quotes. Side effects are spelled out rather than left to inference.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, purpose front-loaded, with every clause earning its place. Slight deduction for the mixed German/English opening and the high information density, which make it slightly harder to parse than a cleaner bilingual split would be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a state-changing tool with no annotations and no output schema, the description is nearly complete: it covers allowed states, versioning behavior, side effects, and illegal transitions. The only gap is the return value format, which is not described despite there being no output schema to supply it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate, but it never explicitly discusses quoteId, workspaceId, or idempotencyKey. The behavioral narrative implies quoteId identifies the quote to revise, and the parameter names are self-evident, but no parameter-level meaning is added and the description only partially compensates for the total lack of schema prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first clause states a specific verb+resource+scope ('Revidiere eine gesendete/abgelehnte/abgelaufene Offerte' = revise a sent/rejected/expired quote), then sharply differentiates itself from siblings: drafts are edited in place (edit_draft_instead) and accepted/converted quotes are refused. An agent can distinguish this from quotes_update and quotes_convert without inspecting schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use (sent/rejected/expired quotes), explicit when-not-to-use (drafts, accepted or converted quotes), and names the alternative behavior (edit_draft_instead). The two parenthetical codes function as routing instructions, leaving no ambiguity about which states this tool handles.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quotes_sendA
Versende eine Offerte (P8 outbound, draft-gated): one atomic write issuing the gap-free O-number (no ledger posting), minting the single-use e-accept token (only its hash is stored), and moving the quote to sent. Returns the local artifact plus transmitted:false / reason:cloud_tier: the OSS core produces the PDF and the accept link and stops; email/portal delivery is the cloud tier (needs_dispatch_module). The G05 send log records one artifact_created row (list_dispatches). A quote with no lines is refused (no_lines); a non-draft quote is refused (illegal_transition). Idempotent: a replay returns the original token and number, never a second of either.
| Name | Required | Description | Default |
|---|---|---|---|
| quoteId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it is exceptionally transparent. It discloses atomicity, no ledger posting, hash-only token storage, single-use token semantics, idempotent replay behavior, the G05 log side effect, and the transmitted:false / reason:cloud_tier return state. This goes far beyond a generic send operation and directly informs the agent of important behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause earns its place: atomicity, token handling, cloud-tier limitation, log effect, error conditions, and idempotency are all relevant. It is front-loaded with the core purpose. The mixed German/English phrasing and the long single paragraph make it slightly harder to parse than a more structured description, but nothing is wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description is remarkably complete: it explains the return artifact, the transmitted:false/reason:cloud_tier outcome, the log row, rejection reasons, and replay behavior. It does not describe the full output structure or permission requirements, but for a mutating send tool it covers the critical operational context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It adds meaning to idempotencyKey through the replay guarantee (“never a second of either”) and clarifies quote-state constraints. However, it never explicitly defines workspaceId or quoteId formats, and it does not map any of the described behavior to the parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource, “Versende eine Offerte” (send a quote), and adds precise scope: P8 outbound, draft-gated, atomic write. It distinguishes itself from siblings like quotes_accept, quotes_revise, and quotes_convert by describing exactly what this send operation does and does not do (no ledger posting, no cloud delivery).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-not conditions: a quote with no lines is refused (no_lines) and a non-draft quote is refused (illegal_transition). It also explains the OSS-core versus cloud-tier split and the needs_dispatch_module requirement, which tells the agent when full delivery will not happen. It does not explicitly name an alternative sibling tool, but the conditions and limitations are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
quotes_updateA
Bearbeite eine Offerte im Entwurf: retitle the contact, re-price lines (re-resolved and re-snapshotted exactly as create), or set validUntil/intro/outro. Draft-only: A10 refuses a patch on an issued or later quote (document_immutable), which is the tax/total freeze at work. A validity in the past is refused (validity_in_past).
| Name | Required | Description | Default |
|---|---|---|---|
| patch | Yes | ||
| quoteId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses that line re-pricing re-resolves and re-snapshots exactly as create, and explains two refusal conditions: document_immutable (tax/total freeze) and validity_in_past. These are non-obvious behaviors. However, it omits idempotency behavior (despite an idempotencyKey parameter) and whether patch is a partial or full replacement, leaving gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the main purpose. It mixes German and English, which is slightly awkward but not confusing. The key details are packed into a single sentence, and the behavioral notes are placed after the primary actions. No wasted words, but the language mixing detracts slightly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex update tool with a nested patch object and no annotations, the description is incomplete. It covers the main constraints but misses idempotency behavior, whether patch is partial, return value expectations, and general error handling beyond the two specific refusals. Given the tool's complexity and lack of annotation coverage, more detail is needed for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions contact, lines, and validUntil/intro/outro, but does not explain currency, idempotencyKey, or the structure of the lines array (e.g., quantityMilli, unitPriceMinor). It also fails to clarify whether patch is partial, which is critical. The description covers only a subset of parameters and provides limited semantic depth.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool edits a quote in draft (Bearbeite eine Offerte im Entwurf) and lists specific actions: retitle contact, re-price lines, set validUntil/intro/outro. It also explicitly limits to draft-only, distinguishing it from issued-quote operations. This makes the purpose unambiguous and separates it from siblings like quotes_create or quotes_send.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states the tool is draft-only and explains that A10 refuses patches on issued quotes, giving a clear condition for use. However, it does not name an alternative tool for issued quotes (e.g., quotes_revise) nor explicitly say 'use X for issued quotes', so while the context is clear, the guidance could be more direct.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rate_card_endA
Beende einen Tarif ohne Nachfolger (a client override lapses): writes valid_to on an open card. An already-ended card refuses with rate_card_already_ended; validTo must be after validFrom.
| Name | Required | Description | Default |
|---|---|---|---|
| validTo | Yes | ||
| rateCardId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses that it writes valid_to on an open card, that an already-ended card is rejected with a specific error (rate_card_already_ended), and that validTo must be after validFrom. These are important constraints and error behaviors. It does not mention return values or idempotency, but the key behavioral traits are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, two sentences, with the action and condition front-loaded. There is no redundant text. It efficiently communicates the purpose, the error condition, and a validation rule in a compact form.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has four parameters, no output schema, and no annotations. The description covers the core behavior, an error condition, and a validation constraint. It does not mention what the tool returns on success or how idempotencyKey behaves, which would be helpful given the lack of output schema. However, for a relatively simple operation, the description provides enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the validTo parameter by stating it must be after validFrom, adding meaning beyond the schema. However, workspaceId, rateCardId, and idempotencyKey are not described. While the parameter names are somewhat self-explanatory, with zero schema descriptions more elaboration on each parameter would improve clarity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (beende = end), the resource (Tarif/rate card), and the specific condition (ohne Nachfolger, i.e., without a successor). It distinguishes itself from sibling tools like rate_card_upsert by specifying it ends a card rather than creates or updates it. The mention of writing valid_to on an open card makes the action unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by stating 'ohne Nachfolger' (without successor), which signals when this tool applies. However, it does not explicitly compare to alternatives like rate_card_upsert or rate_card_list, nor does it state when not to use it. The condition is implied but not contrasted with other rate card operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rate_card_listARead-only
List the rate cards (P5): every version with scope, scopeRef, rateMinor, currency and validity, optionally one scope or only the cards active at a day.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No | ||
| activeAt | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes the safety profile, and the description adds useful behavioral context by stating that every version is returned along with the listed fields and optional filtering. However, it does not disclose ordering, pagination, default scope behavior, or how 'active at a day' is interpreted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence front-loads the core operation and adds only relevant details about returned fields and filters. The unexplained 'P5' label is a minor clarity deduction, just enough to prevent a top rating.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list operation, the description is reasonably complete: it names the resource, the returned attributes, and the two optional filters and relies on the readOnlyHint annotation for safety. It lacks pagination and ordering guidance, but no output schema exists and the rest of the context is adequately covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of explaining the parameters, and it largely does: 'optionally one scope' maps directly to the scope parameter and 'active at a day' maps to activeAt. The required workspaceId is not explained, but its meaning is self-evident from its name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a concrete verb and resource, 'List the rate cards (P5)', and enumerates the returned fields (scope, scopeRef, rateMinor, currency, validity) plus optional filters. This makes the tool easy to distinguish from write-oriented siblings like rate_card_upsert and rate_card_end, despite a cryptic 'P5' parenthetical.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to choose this tool over alternatives, no exclusions, and no mention of related list or rate-card operations. The optional filters indicate what the tool can do, but not the conditions under which it should be preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rate_card_upsertA
Lege einen Tarif an (rate card): scope default, employee, project or client (scoped cards name their scopeRef), integer-Rappen rateMinor, valid from a day. costRateMinor is the optional INTERNAL cost rate (same currency) that values the B03 Projekterfolg cost basis; entries snapshot it at capture like the bill rate, and without one the cost basis degrades honestly (basisDegraded). A new card for the same scope VERSIONS the previous one: the open predecessor is end-dated at the new validFrom and its rate is never mutated, so snapshotted entries keep their price (OP1). Overlapping validity refuses with rate_card_overlap.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | Yes | ||
| currency | No | ||
| scopeRef | No | ||
| rateMinor | Yes | ||
| validFrom | Yes | ||
| workspaceId | Yes | ||
| costRateMinor | No | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses key side effects: end-dates the predecessor, never mutates the existing rate, snapshots cost rate at capture, and degrades the cost basis when absent (basisDegraded). It also mentions the error condition rate_card_overlap. This is strong transparency, though it omits details like authorization or currency defaults.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a dense paragraph without fluff. It front-loads the primary action and then details behavior and edge cases. Each sentence adds meaning; it is not overly long for the complexity it conveys.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
It covers the essential operational details: scope types, rate and cost rate semantics, versioning rules, and overlap error. It lacks an explicit statement about return values (no output schema) and prerequisites like workspace existence, but these are minor given the richness of the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters. It does for scope, scopeRef, rateMinor (integer in Rappen), validFrom (a day), and costRateMinor (optional internal cost rate). It does not explain currency, workspaceId, or idempotencyKey, but those are common and the main domain-specific parameters are covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action (create a rate card), enumerates the scope types (default, employee, project, client), and names the key fields (rateMinor, validFrom). It distinguishes from siblings like rate_card_end and rate_card_list by focusing on the creation and versioning behavior, making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the versioning behavior: a new card for the same scope end-dates the predecessor, and overlapping validity is rejected with rate_card_overlap. This implicitly guides when to use this tool (to add a new rate or version) and when to avoid (overlapping). It does not explicitly name alternatives like rate_card_end, but the context is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
receipt_recordA
AGENT-FIRST. Erfasse einen Wareneingang (voll oder teilweise) gegen eine gesendete Bestellung: writes a goods_receipt + lines and, for every stock-tracked line, mints ONE D01 receipt movement through stock.move (OP2, D02 never writes stock_movement itself), storing each stock_movement_id for 1:1 traceability. Each po_line.received_qty rises; when every line is fully received the PO advances sent -> received. The unit cost handed to D01 is the line`s CHF BASE cost (H-FX). Atomic and idempotent: over-receipt (qty over the open order qty) is PRE-CHECKED and refused with over_receipt (a refusal writes zero rows and moves zero stock), a D01 refusal rolls the whole receipt back, and a retry under the same idempotencyKey mints NOTHING a second time. Receiving against a draft/received/closed/cancelled PO is refused (invalid_transition). The intended agent loop is D01 low-stock -> supplier_price_list -> pre-filled po_upsert -> receipt.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| poId | Yes | ||
| actor | No | ||
| lines | No | ||
| locationId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it excels: it discloses side effects (writes goods_receipt + lines, mints ONE D01 movement, raises po_line.received_qty, advances PO sent->received), cost basis (CHF BASE cost via H-FX), atomic and idempotent behavior, over-receipt pre-check refusal with zero writes, D01 rollback, and retry semantics. Nothing is hidden, and there is no annotation contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence carries essential behavior or routing information not present in the schema or annotations. The front-loaded AGENT-FIRST label and the action verb orient the agent immediately, followed by side effects, failure semantics, and the intended agent loop. There is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity tool with no output schema and no annotations, this description is remarkably complete. It covers success outcomes, all key failure modes (over_receipt, invalid_transition, D01 rollback), unit cost rules, idempotency behavior, state transitions, and the interaction with the D01 low-stock loop. Nothing critical for an agent to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds real meaning for poId (only sent POs; invalid_transition for other states), lines/qty (over-receipt pre-checked and refused), and idempotencyKey (same key mints nothing a second time). However, parameters like note, actor, and workspaceId are left entirely undocumented, so compensation is not complete enough for a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb+resource: 'writes a goods_receipt + lines' against a sent PO, including the D01 movement side effect. It also distinguishes itself from stock_move and goods_receipt_create by clarifying that it mints exactly one D01 movement through stock.move (OP2, D02 never writes stock_movement itself). It is far from a tautology and clearly identifies the tool's singular role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended agent loop is explicit: 'D01 low-stock -> supplier_price_list -> pre-filled po_upsert -> receipt,' which clearly tells the agent when to call this tool. It also gives when-not conditions: receiving against draft/received/closed/cancelled POs is refused with invalid_transition. However, it does not explicitly name alternative tools (e.g., goods_receipt_create for staged receipts), so the exclusion/alternative guidance is incomplete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_exchange_rateA
Record a foreign-currency rate (§H-FX). rate is a DECIMAL STRING and is the price of one unit of baseCurrency in the workspace base currency, so EUR/CHF 0.9412 means 1 EUR = 0.9412 CHF. asOf is the date the rate is valid FOR. method names the admissible MWSTV Art. 45 basis (daily | monthly_avg | bank | group). Recording the same rate twice is one row; a different rate under the same date and source is refused, because it may already have priced a posted entry.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | Yes | ||
| rate | Yes | ||
| method | No | ||
| source | No | ||
| provenance | No | ||
| workspaceId | Yes | ||
| baseCurrency | Yes | ||
| quoteCurrency | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well: it discloses idempotent deduplication (same rate twice is one row), conflict refusal (different rate under same date/source is refused), and the rationale (may already have priced a posted entry). This is meaningful behavioral context beyond the tool name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient: a clear lead sentence followed by parameter-specific semantics and a crucial idempotency warning. No sentence is wasted, though it could be slightly better organized with delimiter-separated parameter notes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without annotations or an output schema, the description needs to cover the full call contract, but it does not. It omits explanation of quoteCurrency, source, provenance, and idempotencyKey, and says nothing about return values or prerequisites. For a complex 9-parameter mutation tool, this is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains rate as a decimal string with direction, asOf as the validity date, and method with enumerated MWSTV values. However, several parameters—workspaceId, quoteCurrency, source, provenance, and idempotencyKey—are left undocumented or only implied, leaving gaps for a 9-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and resource: 'Record a foreign-currency rate', then clarifies the critical rate direction (price of one unit of baseCurrency in the workspace base currency) with an example. This clearly distinguishes the tool from read/list siblings like get_exchange_rate and list_exchange_rates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance or alternatives. It never mentions bulk import (import_exchange_rates), reading rates, or setting FX method, so an agent cannot tell when to choose this tool over related siblings. The behavior is described, but usage selection is left entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_expenseA
Capture AND post a supplier bill or expense in one step (the human default): the same input as create_vendor_bill, composed with post_vendor_bill in ONE transaction, so a refused posting leaves no draft behind. Posts the expense or asset net plus 1170/1171 Vorsteuer against 2000 Kreditoren gross, stamping the §H-VAT-TRACE the MWST-Abrechnung reads. Under the Saldo method (MWSTG Art. 37) nothing is separately reclaimed and the expense books gross: the result says so with vorsteuerDeductible:false.
| Name | Required | Description | Default |
|---|---|---|---|
| fxRate | No | ||
| dueDate | No | ||
| taxCode | No | ||
| billDate | Yes | ||
| currency | No | ||
| vendorId | Yes | ||
| projectId | No | ||
| receiptRef | No | ||
| supplyDate | No | ||
| amountMinor | Yes | ||
| workspaceId | Yes | ||
| costCenterId | No | ||
| amountIsGross | No | ||
| idempotencyKey | Yes | ||
| vendorReference | No | ||
| expenseAccountId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It discloses atomicity (refused posting leaves no draft), the accounting posting details (accounts and VAT codes), and the Saldo method behavior (vorsteuerDeductible:false). This is extensive and goes beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is about four sentences, well-structured, with the main action first and supporting details following. It is dense but not bloated, using technical jargon appropriately for the domain.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 16 parameters, 6 required, and no output schema, the description covers the high-level behavior and VAT logic but omits any explanation of required inputs or parameter meanings. It is not complete enough for an agent to construct a valid call without consulting external documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. However, it only says 'the same input as create_vendor_bill' without describing any of the 16 parameters (e.g., amountMinor, taxCode, idempotencyKey). This provides almost no semantic guidance for the agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Capture AND post a supplier bill or expense in one step' and explicitly differentiates it from create_vendor_bill and post_vendor_bill by composing them in one transaction. This makes the purpose specific and distinguishable from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It calls itself 'the human default', implying it is the standard path, but it does not explicitly state when to use the alternatives (e.g., create_vendor_bill alone for a draft workflow). The context is clear but exclusions are not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_incoming_creditA
Register an incoming bank credit in the Abgleich queue and score it against the open invoices: manual entry today, and the seam A20's camt import will feed later. bankAccountId is the A19 Bankkonto the money landed on (needs_bank_account if unregistered), amountMinor is integer Rappen, valueDate the credit's value date, reference whatever the advice carries (a QRR or Creditor Reference is validated; free text falls to the heuristic lane). Writes the queue row and NOTHING else: no posting, no ledger effect. A bankTxnId already registered returns the existing row (a re-imported statement duplicates nothing), and a repeat call under the same idempotencyKey replays byte-identically.
| Name | Required | Description | Default |
|---|---|---|---|
| currency | No | ||
| bankTxnId | No | ||
| payerName | No | ||
| reference | No | ||
| valueDate | Yes | ||
| amountMinor | Yes | ||
| workspaceId | Yes | ||
| bankAccountId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden and it handles it well: explicit side-effect disclosure ('Writes the queue row and NOTHING else'), idempotency behavior for repeated idempotencyKey, de-duplication via bankTxnId, and validation lanes for references. This goes beyond a generic write tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The main purpose is front-loaded in the first sentence and each sentence adds substance—queue side effects, parameter units, validation, idempotency. The references to A19/A20 are slightly insider-heavy and could be trimmed, but the paragraph remains efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, no-annotation, no-output-schema tool, the description covers the critical invocation facts: side effects, idempotency, units, and error condition for an unregistered bank account. It does not describe the general return value or the other three parameters, so it is very good but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero descriptions, so the prose compensates by explaining bankAccountId, amountMinor, valueDate, reference, and idempotencyKey semantics. However, workspaceId, currency, and payerName are left unexplained, leaving a small but real gap for an agent trying to construct a correct call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Register an incoming bank credit in the Abgleich queue and score it against the open invoices.' It also differentiates from nearby tools by drawing the boundary around queue-only behavior, 'Writes the queue row and NOTHING else: no posting, no ledger effect,' and mentions the future camt-import seam.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear operational context: manual entry now, with A20's camt import later, and explains that amountMinor is Rappen and bankAccountId must be a registered A19 bank account. It does not explicitly name alternatives/siblings or state when not to use the tool, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_paymentA
Record an incoming or outgoing payment and post its balanced entry (bucht eine Zahlung). Allocates across one or several open items with partial amounts, Skonto, and an Ausbuchung of a small residual; any remainder is parked as a Guthaben for its counterparty. Requires intent='post_payment', because money never moves as a side effect. CONSEQUENCE: Books a payment against the ledger and settles the allocated open items.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | ||
| fxRate | No | ||
| intent | Yes | ||
| source | No | ||
| currency | No | ||
| direction | Yes | ||
| reference | No | ||
| allocations | No | ||
| amountMinor | Yes | ||
| workspaceId | Yes | ||
| bankAccountId | Yes | ||
| counterpartyId | No | ||
| idempotencyKey | Yes | ||
| onAccountMinor | No | ||
| counterpartyKind | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing side effects. It clearly warns: 'CONSEQUENCE: Books a payment against the ledger and settles the allocated open items' and explains how remainders are parked as a Guthaben. This goes beyond a generic mutation hint by specifying what exactly changes. It lacks details on idempotency, reversibility, or error cases, but for a tool with no annotation safety profile, the disclosed consequence is substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with dense, purposeful content. It front-loads the primary action and smoothly incorporates the allocation behavior, the intent requirement, and the consequence. The capitalization of 'CONSEQUENCE' adds a slight stylistic noise but doesn't introduce waste. No sentence is redundant; even the consequence reiteration reinforces the safety-critical side-effect clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 15 parameters, no annotations, and no output schema, the description must provide substantial guidance to be usable. It explains the workflow and main effects, but still leaves open critical call details: how to construct allocations, what idempotencyKey is for, how direction/currency/fxRate interact, and what the response contains. The description is complete enough for a domain-savvy agent to understand the tool's role, but not enough to confidently fill every required parameter without guessing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaningful semantics for key parameters: intent must be 'post_payment', allocations involve partial amounts/Skonto/Ausbuchung, and any remainder becomes a Guthaben. However, many of the 15 parameters (fxRate, source, reference, onAccountMinor, counterpartyKind, etc.) receive no explanation in either the schema or the description, leaving the agent to infer their roles. The description covers the core payment-allocation semantics but not the full parameter surface.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Record an incoming or outgoing payment and post its balanced entry.' It also distinguishes itself from related operations like allocate_payment or preview_payment by clarifying that this tool actually books the payment rather than merely preparing or allocating it. The mention of 'bucht eine Zahlung' reinforces the concrete posting behavior, leaving little ambiguity about the tool's role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly requires intent='post_payment' and explains the rationale: 'money never moves as a side effect.' This tells the agent when to invoke this tool (when an actual posting is desired, not just a preview or allocation). It does not name specific sibling alternatives, but the intent requirement and consequence language provide sufficient context to distinguish this from related payment tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
refresh_plugin_compatB
Prüfe die Kompatibilität neu (US-G02.4): compare the plugins compat_range against the current core version with standard semver range semantics and persist the verdict. A range that no longer matches flips status to incompatible and sweeps the plugins capabilities (its manifest and granted permissions retained); a range that matches again flips it back to installed and re-registers. Never a crash, only the badge (P9).
| Name | Required | Description | Default |
|---|---|---|---|
| pluginId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden well: it discloses persistence, status flips, capability sweeping, retention of manifest/permissions, re-registration, and non-crashing error behavior. It still omits auth/permission needs and idempotency semantics, so it is not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded and the side-effect description is dense but useful. However, the requirement tag 'US-G02.4' and the cryptic 'only the badge (P9)' add noise an agent cannot act on, and the mixed German/English wording reduces clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the tool's behavioral outcomes well for a mutation with no annotations and no output schema. But it leaves the required parameters undocumented and does not indicate what a successful response contains or what conditions must hold before the check can run.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage across all three required parameters, and the description does not explain workspaceId, pluginId, or idempotencyKey. In particular, idempotencyKey's role is entirely undocumented, which is a serious gap for a required parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: compare the plugin's compat_range against the current core version using semver semantics and persist the verdict. The focus on compatibility re-checking clearly distinguishes it from the many plugin lifecycle and registry tools in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 're-check' wording implies use after a core version change, and the status-transition explanation implies when the result matters. However, the description never explicitly says when to call this instead of install_plugin, list_plugins, or get_plugin, and offers no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reject_drafted_actionA
Reject a drafted agent action from the inbox: drop it unexecuted (idempotent on actionId; an already-executed action refuses). reason is the human's optional sentence on why, stored on the proposal and shown in the trace beside the drafting call, so the next session reads it. Owner-gated (manage_agent_dial).
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| actionId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full disclosure. It states idempotency, refusal of executed actions, side effects (reason stored on proposal and shown in trace), and owner-gating with permission manage_agent_dial. This goes beyond typical descriptions and fully informs the agent of behavioral nuances.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary action and packs idempotency, refusal, reason handling, and access control into a compact paragraph. Every sentence adds value without fluff, though it could be split for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers key behavioral aspects: idempotency, refusal, side effects, and authorization. It does not mention the return value, but that is not required without an output schema. Overall, it provides sufficient context for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains 'reason' thoroughly (optional sentence, stored, shown in trace) but does not explicitly define 'actionId' or 'workspaceId'. While these are somewhat self-evident from context, the description only partially covers the parameter space.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Reject a drafted agent action from the inbox' with the specific outcome 'drop it unexecuted'. It distinguishes from siblings like approve_drafted_action and list_drafted_actions by focusing on the rejection semantics, including idempotency and refusal of already-executed actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (when a drafted action should be dropped unexecuted) and notes a key constraint (already-executed actions refuse). It doesn't explicitly name alternatives like approve_drafted_action, but the sibling list and the clear purpose make the usage context unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reopen_monthA
Reopen a soft-closed month (a hard-sealed month refuses). CONSEQUENCE: Reopens a soft-closed month, so postings into a month a human had declared done are accepted again.
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are supplied, so the description carries the full burden. It discloses the postcondition (postings are accepted again) and the refusal condition for hard-sealed months. It doesn't mention side effects, reversibility, or idempotency, but the core behavior is clearly transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, with the action and constraint front-loaded. The second sentence repeats 'Reopens a soft-closed month' before adding the consequence, so a small redundancy prevents a perfect score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers what happens and the key refusal case. Missing parameter guidance, period format, and error/response behavior leave clear gaps, making it adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for three required parameters. The description only clarifies that 'period' refers to a month in soft-closed state; workspaceId and idempotencyKey semantics are entirely absent. It does not compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Reopen') and resource ('soft-closed month') and immediately disambiguates by noting that a hard-sealed month refuses. This clearly separates it from sibling tools like close_month, lock_period, and unlock_period even without reading their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly scopes the action to soft-closed months and excludes hard-sealed months with 'refuses'. It gives clear when-to-use context but does not name an alternative tool for hard-sealed months, so it falls short of full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reports_deleteA
Bericht löschen (F01): remove a definition and its run history. Refused with retention_locked while any run is retention-linked into E00 and still locked there (E00 owns the OR 958f lock); retained E00 documents are untouched. Deleting also deactivates any schedule, so a deleted report can never fire again. Idempotent on idempotencyKey. Gated on A24 reports.write.
| Name | Required | Description | Default |
|---|---|---|---|
| reportId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so richly: it discloses destructive scope, run-history removal, schedule deactivation, idempotency behavior, and a lock-based refusal condition. No annotation contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense and front-loaded: the main action appears first, followed by constraints and side effects. Every clause adds meaningful information without filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive delete operation, the description is complete: it covers destructive scope, refusal conditions, schedule side effects, idempotency, and permission requirements. The parameters are simple and there is no output schema, so no essential detail is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains idempotencyKey behavior and implicitly maps reportId to the report definition, but does not clarify workspaceId or how to construct each parameter beyond the names. Partial compensation only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the verb 'remove' and the resource: 'a definition and its run history,' making it a deletion operation for report definitions. The mention that deleting deactivates any schedule further differentiates it from report update/duplicate/run siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear deletion use case, a precise refusal condition (retention_locked while runs are locked in E00), and the permission gate (A24 reports.write). It does not explicitly name alternative tools, but the deletion context and side effects make when to use it clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reports_duplicateA
Bericht duplizieren (F01): copy a definition with a "(Kopie)" suffix and WITHOUT its schedule or recipients, so a copy never silently starts mailing anyone (P8 spirit). Idempotent on idempotencyKey. Gated on A24 reports.write.
| Name | Required | Description | Default |
|---|---|---|---|
| reportId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the transparency burden and does substantial work: it discloses the no-schedule/no-recipients safeguard, idempotency on idempotencyKey, and the reports.write permission gate. It does not describe the return payload, but the critical behavioral traits are exposed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: two dense sentences cover the action, the safety boundary, idempotency, and permissions. Internal codes like F01, P8, and A24 add some jargon, but they do not undermine readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter duplicate operation, the essential invocation details and side-effect boundaries are covered. However, with no annotations and no output schema, the lack of any return-value or post-invocation behavior leaves a notable gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only adds explicit semantics for idempotencyKey (idempotent key). reportId and workspaceId are left to inference from their names and the phrase 'copy a definition', which is reasonable but not fully documented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action—duplicate a report definition—with concrete behavior (adds a '(Kopie)' suffix) and a key distinguishing constraint (does NOT copy schedule or recipients). This clearly differentiates it from sibling tools like reports_save, reports_update, and reports_schedule.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: when a copy should be created without silently inheriting mailing behavior. It also states idempotency and permission constraints. However, it does not explicitly name alternatives or give when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reports_listARead-only
Berichte anzeigen (F01): the saved-report list (P5), tenant-scoped, newest first, with each definition source, format, schedule, recipients (draft), delivery state and last-run timestamp. Gated on A24 reports.read.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the agent knows it is safe. The description adds that the list is tenant-scoped, newest first, and gated on A24 reports.read, which is useful beyond the annotation. It does not mention pagination, full availability semantics, or what happens with drafts, but the extra detail about the gate and ordering earns a middle score.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One dense sentence with no filler, front-loading the purpose and then packing the list of returned fields and the access gate. Every part adds information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is strong on what the tool returns and its access gate, but with no output schema and one undocumented parameter, an agent still lacks the meaning of the only required input and pagination details. For a simple read-list tool this is acceptable, but it is not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the entire burden for explaining the sole parameter, workspaceId. The description does not mention workspaceId at all beyond the tenant-scoped phrase, leaving the agent to guess that the parameter is the required tenant identifier. This is a significant gap for a one-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Berichte anzeigen' / show reports) plus the resource ('saved-report list'), and it states the tenant scope, ordering, and the fields included. It reads as distinct from sibling reports_runs, reports_save, reports_schedule, and reports_preview because it is specifically the list of saved report definitions, not executions or previews.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool (when you need the saved-report list with definition metadata) by naming the resource and the fields it returns. It does not explicitly state when not to use it or point to alternatives such as reports_runs for run history or reports_preview for output. With a large sibling set of report tools, explicit routing would have helped.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reports_previewARead-only
Vorschau (F01): read-only ad-hoc compute of a source read model with filters + columns for the builder live preview and agent exploration (P5). Persists nothing. Returns the projected rows (up to limit, default 50), the resolved columns and the full row count. Asserts the source own read gates (a preview never reads past the caller RBAC). Unknown source returns unknown_source; a filter on an unpublished field returns invalid_filter_field.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| source | Yes | ||
| columns | Yes | ||
| filters | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint: true, but the description adds crucial behavioral context: persists nothing, asserts the source's read gates, never reads past caller RBAC, and returns error codes for unknown source and invalid filter fields. This goes well beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-organized, front-loading the core purpose and then covering persistence, return values, RBAC, and error cases. Every sentence adds value, though it could be tightened slightly by trimming the parenthetical identifiers (F01, P5) that are not essential for agent understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a preview tool, the description covers purpose, persistence, RBAC, return values, and error scenarios. It does not specify the exact JSON response shape, but since there is no output schema, this is a minor gap. The behavior is sufficiently described for an agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must clarify parameters. It mentions 'filters + columns' and explains the limit parameter with its default of 50. However, it does not detail the structure of the filters or columns arrays (e.g., field names, operators), nor does it clarify workspaceId or source beyond being required. It partially compensates but leaves ambiguity in parameter formats.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool computes a read-only preview of a source read model with filters and columns, for builder live preview and agent exploration. It distinguishes itself from report execution/saving tools by emphasizing it persists nothing and is ad-hoc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly scopes usage to 'builder live preview and agent exploration (P5)', implying it is for interactive exploration rather than final reporting. It also notes it respects RBAC, but does not explicitly name alternative tools like reports_run for persisted reports, so a clear when-not clause is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reports_runA
Bericht ausführen (F01): compute the source read model with the stored filters (P5), project the stored columns, render CSV (RFC-4180, UTF-8 BOM, machine-neutral integer Rappen) or PDF (report header + rows, de-CH money) as a LOCAL artifact (OP4), append a report_runs row and return the content. An empty result set is a valid ok run (row_count 0, header-only CSV). With retain:true an accounting-record source additionally links the artifact into E00 (F01 never computes a retention date; E00 owns OR 958f). Re-running the same idempotencyKey returns the original artifact. Asserts the source own read gates too. Gated on A24 reports.run.
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | ||
| retain | No | ||
| reportId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it delivers: it states side effects (append report_runs row, link into E00 when retain:true), idempotency behavior, empty-result handling, format specifics, local artifact creation, and permission gating ('Gated on A24 reports.run'). This is far richer than a generic 'run report' statement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the core verb and pipeline, then moves through edge cases and constraints. Every clause adds a distinct behavioral fact with minimal filler. The heavy use of internal codes (P5, OP4, E00, OR 958f) makes it less scannable, but the text is still appropriately sized for the amount of semantics packed in.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero annotations, zero schema descriptions, and no output schema, the description is remarkably complete: it covers input semantics, output formats, side effects, idempotency, empty runs, retention linkage, and authorization. The main gap is that it does not explicitly state the accepted format parameter values or the exact response envelope structure, but the overall contract is well specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does meaningfully: it explains format semantics (CSV RFC-4180/UTF-8 BOM vs PDF de-CH money), retain behavior (E00 linkage for accounting-record sources), and idempotencyKey behavior (re-run returns original artifact). reportId and workspaceId remain implicit, but their roles are inferable from the report-run context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is explicit: 'Bericht ausführen' plus the detailed pipeline (source read model -> project stored columns -> render CSV/PDF -> append report_runs row -> return content) leaves no doubt about what the tool does. This clearly distinguishes it from sibling tools like reports_preview, reports_save, and reports_schedule, which serve different stages of the report lifecycle.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: it executes a saved report and returns a CSV/PDF artifact. It also gives valuable behavioral conditions such as idempotent re-runs, empty-result validity, and retain:true linkage. However, it never explicitly contrasts this tool with close siblings like reports_preview or reports_schedule, nor does it state when not to use it. The usage guidance is present but implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reports_runsARead-only
Laufverlauf (F01): the append-only run history for one report (P5), newest first, each run status (ok|failed), row count, format, artifact ref, definition-version hash and E00 document link. savedViewId applies a saved view (G00) over the run history: its stored status filter merges underneath an explicit status named here. Gated on A24 reports.read.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| reportId | Yes | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true. The description adds meaningful behavioral context beyond that: the data is append-only, results are newest first, and savedViewId's stored status filter merges underneath an explicit status parameter. It does not cover pagination or limits, but the key behavioral nuances are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and efficient: it front-loads the core purpose and return fields, then explains savedViewId behavior, then notes the permission gate. Every clause carries useful information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description enumerates the return fields (status, row count, format, artifact ref, hash, document link), which is valuable. It also states the required permission. It lacks pagination or response-size details, but for a run-history list tool the essential context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains status and savedViewId semantics, including how the saved view filter merges with an explicit status. However, it does not explain workspaceId or reportId beyond implying reportId selects the report whose history is returned. This is partial compensation for zero schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as returning the append-only run history for a single report, with newest-first ordering and a specific set of fields (status, row count, format, artifact ref, hash, document link). This distinguishes it from sibling tools like reports_run, reports_sources, and reports_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: this is for one report's run history, and it explains how savedViewId filters interact with an explicit status filter. It does not explicitly name alternatives or state when not to use it, but the scope and filtering behavior are clearly implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reports_saveA
Bericht speichern (F01): save a report definition over a REPORT_SOURCES read model (source + filters + columns + format csv|pdf). Validates the source (unknown_source), that every filter names a published field with a type-valid operator (invalid_filter_field), and that columns is a non-empty subset of the published columns, base OR cf: custom fields (columns_empty). Posts nothing (P3); renders no artifact (that is reports_run). Idempotent on idempotencyKey. Gated on A24 reports.write.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| format | No | ||
| source | Yes | ||
| columns | Yes | ||
| filters | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool is idempotent on idempotencyKey, which is a critical behavioral trait for retries, and mentions it is gated on A24 reports.write permission, which is important for authorization. It also explicitly states 'Posts nothing (P3)' and 'renders no artifact', which clarifies side effects. These go beyond what a typical schema provides, earning a 4; it could be a 5 if it also disclosed return behavior or error handling beyond the listed validation errors, but those are partially covered by the error names.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense paragraph, but the information is packed efficiently without fluff. It front-loads the primary purpose ('Bericht speichern (F01): save a report definition') then systematically covers validation, side effects, idempotency, and permissions. The sentence about 'Posts nothing (P3)' is slightly cryptic but not wasteful. It could be improved by breaking into bullet points, but given its concise nature, a 4 is appropriate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (7 parameters, validation logic, idempotency), the description covers the core purpose, validation rules, side effects, and permissions, which is quite complete. However, it lacks details on the format parameter's allowed values (only mentions 'csv|pdf' in passing), how to construct filters and columns (no examples), and what the response contains (no output schema). The absence of parameter descriptions in the schema makes this incomplete for an agent to call it without further information, so a 3 is fitting.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the schema provides no descriptions for the 7 parameters. The description does add meaning: it explains that source must be over a REPORT_SOURCES read model, filters must reference published fields with type-valid operators, and columns must be a non-empty subset of published columns (base or custom fields). However, it does not explain important parameters like name, format (except 'csv|pdf' inline), workspaceId, and idempotencyKey semantics. With 0% coverage, the description should compensate more thoroughly; it partially does but leaves gaps for required parameters like idempotencyKey and workspaceId, so a 2 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: save a report definition, and specifies the resource (REPORT_SOURCES read model) and key components (source, filters, columns, format). It distinguishes from reports_run by explicitly stating it does not render an artifact. However, it does not explicitly contrast with other report-related siblings (e.g., reports_update, reports_duplicate, reports_preview), which could cause confusion for an agent selecting among them, though the 'save' action is distinct enough for basic differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: it is for saving a report definition over a specific read model, and explicitly notes that rendering is done by reports_run, which helps route to the right sibling. However, it does not state when to use this tool versus reports_update or reports_create (if exists) for modifying existing reports, nor does it provide exclusions for other sibling tools in the reports family. The context is good but lacks explicit alternatives beyond reports_run.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reports_scheduleA
Zeitplan setzen (F01): store a validated cron-subset schedule (freq daily|weekly|monthly + at HH:mm, weekly weekday, monthly dayOfMonth) on a saved report; an unsupported expression is refused (invalid_schedule). Recipients are stored draft-gated (P8); in the OSS core there is no delivery transport, so delivery stays inactive (reason cloud_tier) and the LOCAL run always works. schedule:null clears the schedule and keeps the report. Idempotent on idempotencyKey. Gated on A24 reports.write.
| Name | Required | Description | Default |
|---|---|---|---|
| reportId | Yes | ||
| schedule | No | ||
| recipients | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With zero annotations, the description carries the full burden and excels: it discloses error behavior ('an unsupported expression is refused (invalid_schedule)'), idempotency ('Idempotent on idempotencyKey'), permission gating ('Gated on A24 reports.write'), draft-gating of recipients (P8), and the critical runtime limitation that 'in the OSS core there is no delivery transport, so delivery stays inactive (reason cloud_tier) and the LOCAL run always works.' This is exactly the behavioral context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: core action first, then validation/error behavior, recipients, delivery limitation, null clearing, idempotency, and auth gating. It is dense but compact — roughly four sentences covering seven distinct facts with no repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-param tool with a nested schedule object, 0% schema coverage, no annotations, and no output schema, the description is remarkably complete: format, errors, gating, idempotency, delivery behavior, and clearing semantics are all covered. The only gaps are the return value on success (no output schema exists to cover this) and what happens when the optional schedule parameter is omitted entirely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate — and it largely does. It adds meaning to the schedule object (validated cron-subset with freq/at/weekday/dayOfMonth), explains null semantics for clearing, clarifies the idempotencyKey role ('Idempotent on idempotencyKey'), and notes recipients are draft-gated. Only workspaceId and reportId are left implicit, but those are self-evident from the schema names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: 'store a validated cron-subset schedule ... on a saved report.' It explains the exact schedule format (freq daily|weekly|monthly, at HH:mm, weekday, dayOfMonth) and what schedule:null does, which clearly distinguishes it from siblings like reports_run (executes a report) and reports_save/reports_update (create/modify the report itself).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it — to set, update, or clear a schedule on a saved report — and clarifies the clearing use case ('schedule:null clears the schedule and keeps the report'). However, it never names alternatives or exclusion conditions, such as when to use reports_run instead or how this differs from create_recurring_schedule in the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reports_sourcesARead-only
Datenquellen (F01): the REPORT_SOURCES registry as a read model, each source id, title, entity kind, accounting-record flag, module availability, and its published columns, the union of the source own base fields plus any cf: custom fields defined on its entity kind (OP7), so a custom field appears the moment G00 defines it. Gated on A24 reports.read.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already marks it read-only, but the description adds valuable behavioral context: it is 'a read model' (reinforcing non-mutation), and it discloses that custom fields (cf:) appear dynamically once defined (OP7), and that access is gated on A24 reports.read. This goes beyond the annotation and informs the agent about permission requirements and dynamic data updates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence that packs multiple concepts (registry, fields, dynamic custom fields, permission gate) without clear breaks. It is not front-loaded with the most critical information; starting with 'Datenquellen (F01)' may confuse. While concise in length, the structure is unwieldy and could be split into clearer sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, the description provides a solid list of returned fields (source id, title, entity kind, accounting-record flag, module availability, published columns) and explains the dynamic inclusion of custom fields. It covers the main behavioral aspects, but misses explaining the workspaceId parameter or any pagination/default behavior. For a simple read with one parameter, it is fairly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description does not mention the workspaceId parameter at all. With only one parameter and no description, the agent is left to infer its meaning, though it is a common identifier across many tools. The description fails to compensate for the schema's lack of documentation, so this dimension scores low.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a read model for the REPORT_SOURCES registry, listing the exact fields returned (source id, title, entity kind, accounting-record flag, module availability, published columns). It distinguishes itself from siblings like reports_runs (execution) and reports_list (listing reports) by specifying its read-only nature and content. The verb is implicit but the resource and scope are explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for retrieving data source definitions needed for report building, and mentions a permission gate (A24 reports.read), but does not explicitly state when to use this tool versus alternatives like reports_preview or reports_runs. No explicit exclusions or alternative routing is provided, leaving the agent to infer the context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reports_updateA
Bericht bearbeiten (F01): patch name/filters/columns/format on a saved report. Re-validates the definition against its source; changing source to one that no longer exists is refused (unknown_source). Editing a definition never mutates past report_runs artifacts (history is append-only). Idempotent on idempotencyKey. Gated on A24 reports.write.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | No | ||
| reportId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full behavioral burden and does so thoroughly: it discloses re-validation against the source, the unknown_source refusal, the append-only history guarantee, idempotency on idempotencyKey, and the A24 reports.write permission gate. These are exactly the behavioral traits an agent needs to anticipate side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded with a clear verb and resource, and every sentence earns its place by adding a distinct behavioral fact: what is patched, validation, history append-only semantics, idempotency, and permission gating. The description is dense but compact, with no repetition or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations or output schema, the definition covers core execution semantics: what fields can be patched, the source validation rule, the no-mutation guarantee, idempotency, and the required permission. It does not describe the success response, the full nested patch structure, or merge/replace behavior, so it is strong but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does by adding real meaning to the opaque patch object, listing its intended contents (name/filters/columns/format), and explaining idempotencyKey behavior explicitly if idempotencyKey. workspaceId and reportId are left to their self-explanatory names, so the description covers the semantically ambiguous parameters well without documenting every parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('patch') on a clear resource ('a saved report') and names the exact fields being modified: name/filters/columns/format. This semantic content immediately distinguishes reports_update from siblings such as reports_save, reports_delete, reports_preview, and reports_run, so an agent can select this tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'on a saved report' gives implied usage context for editing existing report definitions, but the description never explicitly names alternatives or says when not to use this tool. There is no explicit routing among the many reports_* siblings, so the usage guidance is somewhat implicit rather than actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
requisition_approveA
Approve a pending requisition. Completes the named approval task (or the sole open one when taskId is omitted) and, once every required step is satisfied, moves the requisition to approved. Appends an immutable approval event.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | No | ||
| comment | No | ||
| workspaceId | Yes | ||
| requisitionId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses that it completes a specific approval task, falls back to the sole open task when taskId is omitted, moves the requisition to approved only after all required steps are satisfied, and appends an immutable approval event. It does not cover permissions, reversibility, or error behavior, but the core side effects are clear.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loaded with the core purpose, and every sentence adds useful detail (task resolution, state transition, immutability). No redundant or vague wording exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description gives a solid behavioral overview and explains the approval workflow, but with no annotations, no output schema, and 0% parameter coverage, an agent still lacks crucial information about required parameter semantics (especially idempotencyKey), return values, and preconditions beyond 'pending'. It is adequate but has clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for missing parameter documentation. It only explains taskId behavior; workspaceId, requisitionId, idempotencyKey, and comment are not described at all. Required fields like idempotencyKey are left entirely to inference, which is a significant gap for a mutation tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Approve a pending requisition') and explains the approval workflow, making it clear this tool handles requisition approvals. It is easily distinguished from sibling tools like requisition_submit, requisition_reject, or approve_entry because it names the resource and the state transition.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use the tool: when a pending requisition needs approval and an approval task is to be completed. However, it does not explicitly mention alternatives or exclusions (e.g., use requisition_reject to reject), so it stops short of full when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
requisition_cancelA
Cancel a draft, pending or approved requisition that has no conversion yet (status cancelled). A requisition that already has any conversion is has_conversions and must be closed instead.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| workspaceId | Yes | ||
| requisitionId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It does add a key rule (cannot cancel if conversions exist; must close instead) and indicates the resulting status ('cancelled'). However, it does not disclose reversibility, permissions, idempotency behavior, or side effects, which would be expected for a mutation with no annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero wasted words. The action and primary condition are front-loaded, and the alternative is stated in the second sentence. It is appropriately compact for the information provided.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers when to cancel and the key conversion rule, but with no output schema, no annotations, and 0% parameter coverage, it leaves the agent without guidance on required parameters, idempotency semantics, or expected return behavior. An agent could not reliably invoke this correctly based solely on the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it provides no explanation for workspaceId, requisitionId, idempotencyKey, or reason. The agent gets no additional meaning beyond the parameter names in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (cancel) and resource (requisition), scopes it to draft/pending/approved statuses, and explicitly contrasts it with 'close' for requisitions with conversions. This clearly differentiates the tool from sibling requisition_close and other requisition actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use the tool (no conversion yet) and when not to ('must be closed instead' if has_conversions). The alternative is named directly, leaving no ambiguity about routing between cancel and close.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
requisition_closeA
Close a converted or partially-converted requisition (status closed): the demand is settled and no further conversion follows.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| requisitionId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses the resulting status and that no further conversion follows, but omits side effects on linked purchase orders, reversibility, required permissions, and the role of the idempotencyKey parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. The main action, the required requisition state, and the consequence are all conveyed efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The core purpose is clear enough for a simple close action, but the lack of annotations, no parameter explanations, and no distinction from requisition_cancel leave meaningful gaps. It is minimally viable with notable gaps around idempotency and workflow position.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds no meaning to workspaceId, requisitionId, or idempotencyKey beyond their parameter names. The names are self-explanatory, but the description does not compensate for the missing schema documentation, especially for idempotencyKey's behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Close a converted or partially-converted requisition' and pins the terminal state to 'status closed.' This distinguishes it from siblings like requisition_cancel and requisition_convert_to_po by making the target state and post-condition explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear eligibility context: the requisition must already be converted or partially converted, and the demand is settled with no further conversion to follow. It does not explicitly name alternatives such as requisition_cancel or state when not to use this tool, so it stops short of full exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
requisition_convert_to_poA
Convert selected open quantities of an approved (or partially converted) requisition into a D02 purchase order. lines is a list of { lineId, qtyMilli } where each qtyMilli is a positive whole-unit multiple (1000 = one unit) and must not exceed the line open quantity (over_conversion). supplierContactId overrides the lines preferred supplier (required when the selection has none in common). createAs is draft | sent. The requisition becomes partially_converted while any open quantity remains, else converted, with an immutable conversion link back to the PO.
| Name | Required | Description | Default |
|---|---|---|---|
| lines | Yes | ||
| createAs | No | ||
| workspaceId | Yes | ||
| requisitionId | Yes | ||
| idempotencyKey | Yes | ||
| supplierContactId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It discloses the state transition (partially_converted vs converted), the immutable conversion link back to the PO, validation constraints (positive whole-unit multiples, not exceeding open quantity), and the effect of supplierContactId and createAs. This is a thorough disclosure of side effects and constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: it opens with the core operation, then explains the key parameter semantics, then the state outcome. There is no filler, and the most critical constraints are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of annotations and output schema, the description is remarkably complete: it defines inputs, constraints, and resulting state. The main gap is the unexplained idempotencyKey parameter, which is required but not described. For a transaction-like conversion tool, this is a meaningful omission, though the rest is well-covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does for lines, qtyMilli semantics, over_conversion constraints, supplierContactId, and createAs. However, it does not explain idempotencyKey, workspaceId, or requisitionId, leaving some required parameters without added meaning beyond their names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: converting selected open quantities of an approved (or partially converted) requisition into a D02 purchase order. It also clarifies the scope by specifying it is not a general PO creation but a requisition-driven conversion, distinguishing it from nearby tools like po_upsert, requisition_approve, and requisition_submit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear prerequisites and conditions: the requisition must be approved (or partially converted), only open quantities can be converted, and supplierContactId is required when lines have no common preferred supplier. It does not explicitly name alternative tools or state when not to use it, but the context is clear enough for an agent to select it appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
requisition_getARead-only
Read one requisition: header, lines with open quantities, the full approval-event history, the open approval tasks, and the conversion links to any resulting purchase orders.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| requisitionId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description does not need to restate safety. It adds value by specifying the returned components (header, lines, history, tasks, conversion links), which is useful behavioral context beyond the annotation. It does not mention side effects (correctly, as it's read-only) or error conditions, but given the annotation coverage, the added detail about content is sufficient to justify a moderate score.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the core action ('Read one requisition') and then lists the returned data components with no redundancy. It earns its place with concrete, specific details and avoids filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read operation with readOnlyHint annotation and a simple two-parameter schema, the description is fairly complete. It tells the agent exactly what data will be returned (header, lines, history, tasks, conversion links), which is sufficient to decide whether to call it. Minor missing details like response format or pagination are not critical for a single-get operation, and the safety profile is already covered by annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of explaining the parameters. It does not mention workspaceId or requisitionId at all, relying on their self-explanatory names. While the names are clear, the description provides no additional semantics, such as required format, scope, or relationship to other entities. This falls short of compensating for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the operation ('Read one requisition') and enumerates the specific data it returns: header, lines with open quantities, approval-event history, open approval tasks, and conversion links to purchase orders. This distinguishes it from list or modification tools and from sibling requisition tools. The verb+resource is explicit and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: it is the read/get for a single requisition, contrasting implicitly with requisition_list or requisition_my_pending_approvals. However, it does not explicitly state when to use this versus alternatives, nor does it mention any exclusion conditions or prerequisites. The usage context is clear from the name and content, but no direct guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
requisition_listARead-only
List requisitions with filters: status (one value or an array), requesterId, projectId, costCenterId, a neededBy date range (neededByFrom / neededByTo) and a free-text search q over number and description. Accepts a savedViewId (G00 saved-view seam). Ordered by number, newest first.
| Name | Required | Description | Default |
|---|---|---|---|
| q | No | ||
| status | No | ||
| projectId | No | ||
| neededByTo | No | ||
| requesterId | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| costCenterId | No | ||
| neededByFrom | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes that this is a safe read operation, and the description does not contradict it. The description adds useful behavioral details like ordering and filter semantics, but it does not disclose pagination behavior, response shape, or any rate or data-volume limits, so it provides only modest value beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, with the core purpose stated in the first phrase and the rest as dense, relevant filter details. The parenthetical '(G00 saved-view seam)' is cryptic and adds little for an AI agent, keeping this from a perfect conciseness score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool, the description covers the main invocation concerns: filters, ordering, saved-view support, and the free-text search scope. It does not mention pagination or the shape of returned requisitions, and with no output schema present this is a minor but real gap; still, the essential selection and invocation information is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the burden of explaining parameters. It notably explains that status accepts a single value or an array, q searches over number and description, and neededByFrom/neededByTo form a date range. The only required parameter, workspaceId, is not described, though its necessity is already explicit in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource: 'List requisitions with filters.' It enumerates the filter dimensions and ends with ordering behavior, making the operation's scope unmistakable and distinguishing it from single-record or mutating requisition siblings like requisition_get or requisition_upsert.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly communicates when to use this tool: whenever a filtered, ordered list of requisitions is needed. It does not explicitly call out alternatives or exclusions, such as using requisition_get for a single record, but the list-oriented wording provides strong contextual guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
requisition_my_pending_approvalsARead-only
The open approval tasks the workspace still owes a decision on, each joined to its requisition header (number, requester, needed-by, urgency, estimated total). The approval inbox read. Accepts a savedViewId (G00 saved-view seam).
| Name | Required | Description | Default |
|---|---|---|---|
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description adds meaningful context beyond that: it returns open approval tasks joined to requisition headers and describes the returned attributes. It does not contradict the read-only annotation, though pagination or scope limitations are not disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three concise sentences with the core purpose front-loaded. The phrase 'The approval inbox read' is slightly awkward, and 'G00 saved-view seam' is terse, but overall the description is compact and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers what is returned and hints at saved-view filtering, and the read-only annotation covers the safety profile. However, the required workspaceId is unexplained, savedViewId semantics are vague, and there is no output schema. For a two-parameter read tool this is adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It explains savedViewId as a saved-view seam, but the required workspaceId is never described, and 'G00 saved-view seam' is cryptic. The description does not adequately compensate for the schema's lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (open approval tasks, 'approval inbox') and a read action, plus the exact joined fields returned (number, requester, needed-by, urgency, estimated total). The 'read' framing clearly distinguishes it from sibling actions like requisition_approve and requisition_reject.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this is the tool to consult for pending approvals that still need a decision, but it does not explicitly state when to prefer it over alternatives or name any sibling tools. Usage context is evident enough, yet no exclusions or routing guidance are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
requisition_rejectA
Reject a pending requisition with a reason. Status becomes rejected (terminal), all open approval tasks are closed, and an immutable approval event records the decision.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| taskId | No | ||
| workspaceId | Yes | ||
| requisitionId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses key side effects: status becomes rejected (terminal), open approval tasks are closed, and an immutable approval event records the decision. This is informative, though it omits details like idempotency behavior or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that efficiently conveys the action and key consequences. No wasted words; every clause adds meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters, no annotations, and no output schema, the description covers the primary outcome but leaves out parameter semantics and usage guidance. It also does not mention any prerequisites beyond 'pending' status, which is implied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage and the description does not explain any parameters except implicitly mentioning 'reason'. It does not clarify the purpose of workspaceId, requisitionId, idempotencyKey, or taskId, leaving the agent to infer from names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (reject) and resource (pending requisition) with a clear outcome: status becomes rejected (terminal). This clearly distinguishes it from sibling tools like requisition_approve, requisition_return, or requisition_cancel, which have different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for rejecting pending requisitions but does not explicitly state when to use it versus alternatives like requisition_cancel or requisition_close. It lacks explicit exclusions or conditions that would guide selection among the many requisition-related sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
requisition_returnA
Return a pending requisition to the requester for revision, with a reason. Status returns to draft, open tasks are cancelled, and the requester can edit and re-submit; the prior approval events stay queryable.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | Yes | ||
| taskId | No | ||
| workspaceId | Yes | ||
| requisitionId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full disclosure burden and largely delivers: it states the status moves to draft, open tasks are cancelled (a consequential side effect), and prior approval events remain queryable. It omits permission requirements and precondition/failure cases, but the core side effects an agent needs to weigh before invoking are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler: the first leads with the action and reason, the second packs the consequences (draft status, cancelled tasks, re-submission, retained approval events). It is front-loaded and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating workflow tool with no output schema and no annotations, it discloses the workflow's effects well but leaves material gaps: the meaning of taskId, idempotencyKey semantics, and return value are all undocumented. The agent knows what will happen to the requisition but not what to send in each field or what it will receive back.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for five parameters, yet the description only explains 'reason' ('with a reason'). taskId is non-obvious and entirely unexplained — the cancellation-of-open-tasks hint doesn't clarify whether it targets the return or filters affected tasks — and idempotencyKey semantics are absent. The description does not compensate for the schema's total lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource — 'Return a pending requisition to the requester for revision, with a reason' — and adds the resulting state (back to draft, editable, re-submittable), so intent is unambiguous. It does not explicitly name sibling tools like requisition_reject or requisition_cancel, but the state-transition detail functionally distinguishes it from them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied by 'pending requisition... for revision' — this is the tool for sending a request back rather than rejecting or cancelling it. However, the description never explicitly names alternatives (requisition_reject, requisition_cancel) or states when NOT to use it, so an agent must infer selection criteria from workflow effects.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
requisition_submitA
Submit a draft for approval. The policy evaluator decides the required approvers: a zero-estimate requisition auto-approves (status approved) and any positive estimate becomes pending_approval with an open approval task. Submit of an already-pending or terminal document is invalid_transition; a line-less draft is invalid_line.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| requisitionId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure, and it excels: it reveals the state machine (approved, pending_approval), the existence of an approval task, and two error conditions (invalid_transition, invalid_line). This goes far beyond the generic 'submit' expected, giving the agent concrete expectations about side effects and failure modes. No contradictions with annotations exist.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, each packed with essential information. The first sentence states the purpose, the second explains the policy branching, and the third lists invalidity conditions. There is no filler, and the key action is front-loaded. The grammar is efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The behavioral context is richly described, making the tool's state transitions and error conditions clear. However, the absence of any parameter explanation and the lack of an output schema leave gaps: the agent doesn't know what the response contains (e.g., status, approval task ID) or how the idempotency key must be used. These are notable but secondary to the excellent behavioral coverage, hence a 4.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, and the tool description does not mention workspaceId, requisitionId, or idempotencyKey at all. The agent is left to infer that requisitionId identifies the draft, workspaceId scopes the operation, and idempotencyKey ensures replay safety. None of these are explained, and the description does not compensate for the schema gap. This is a major deficiency for a 3-parameter required call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Submit a draft for approval.' It further distinguishes this from other requisition operations by detailing the policy evaluator's decision outcomes (auto-approve vs. pending_approval), making it clear this is the submission step, not approval/rejection/cancellation. The behavior clauses eliminate ambiguity against sibling tools like requisition_approve or requisition_reject.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when to use this tool (submitting a draft) and even provides conditional behavior (zero-estimate auto-approves, positive estimate goes pending). It also lists invalid scenarios (already-pending, terminal, line-less). However, it does not explicitly name alternative tools or state 'use X instead for…', leaving some inference to the agent. The context is strong enough to be a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
requisition_upsertA
Create or edit a DRAFT requisition (the controlled internal-demand document that opens the procure-to-pay chain). requesterId defaults to the caller; neededBy (ISO date) and urgency (normal | high | critical) are the header; each line carries an optional itemId (else a required description), qtyMilli (> 0, integer milli-units where 1000 = one whole unit), an optional uom, estimatedUnitCostRappen (>= 0), and an optional preferredSupplierId. The header total is the exact integer sum of the line estimates. Editing is refused once the requisition has left draft (invalid_transition). Posts nothing and emits no outward artifact.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | ||
| lines | No | ||
| urgency | No | ||
| currency | No | ||
| neededBy | No | ||
| projectId | No | ||
| description | No | ||
| requesterId | No | ||
| workspaceId | Yes | ||
| costCenterId | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Given no annotations, the description carries the burden of behavioral disclosure. It states that editing is refused on invalid_transition, posts nothing, and emits no outward artifact, which are non-obvious side-effect behaviors. It also mentions the header total is computed as an exact integer sum, but it does not clarify idempotency or concurrency semantics beyond the idempotencyKey parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense paragraph that front-loads the core purpose and state constraint, then efficiently enumerates key fields. Every sentence adds substantive value, with no filler or redundancy. The structure is well-organized, moving from purpose to field explanation to behavioral constraints.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential semantics for a draft upsert but lacks some completeness: it doesn't specify validation rules beyond qtyMilli and cost ranges, nor the response format (though no output schema exists). Given the tool's complexity and the lack of annotations and output schema, more details on error conditions and idempotency behavior would be expected for full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, so the description must compensate, and it does richly. It explains the purpose of requesterId, neededBy, urgency, and line-level fields like itemId, qtyMilli, uom, estimatedUnitCostRappen, and preferredSupplierId. However, it does not cover all 11 parameters (e.g., currency, projectId, costCenterId, description) in detail, but the main business semantics are covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool creates or edits a DRAFT requisition, defines it as a controlled internal-demand document, and distinguishes itself from the many requisition_* siblings (submit, approve, reject, etc.) by focusing on the draft state. It specifies the key fields and constraints, making its purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for draft creation/editing and mentions the draft-only constraint ('Editing is refused once the requisition has left draft'), but it does not explicitly guide when to use this tool versus alternatives like requisition_submit or requisition_convert_to_po. The draft context is clear, but explicit exclusion of other states is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
restore_backupA
Restore a .tillbackup into a brand-new workspace: composes create_workspace, bulk-loads every table preserving ids, and re-verifies balance + referential integrity before commit (all-or-nothing). Agent-staged: pass confirmed:true to proceed, else it returns a plan and writes nothing. Never overwrites an existing workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | ||
| confirmed | No | ||
| idempotencyKey | No | ||
| newWorkspaceName | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility for behavioral disclosure. It reveals crucial traits: staged execution (plan vs. commit), bulk-loading with id preservation, balance and referential integrity re-verification before commit, all-or-nothing atomicity, and a guarantee to never overwrite existing workspaces. This is far more transparent than typical tool descriptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three concise sentences that front-load the core purpose, then add behavioral details and a safety constraint. Every sentence provides critical operational information with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the compound nature (4 params, no annotations, no output schema), the description is nearly complete. It explains the two-phase confirmation, atomicity, verification, and non-overwrite behavior. It could explicitly mention what the plan contains or what error is returned if the workspace name already exists, but these are minor gaps given the strong coverage of the workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% coverage, so the description must compensate for parameter meaning. It clarifies 'source' as the .tillbackup file, 'newWorkspaceName' as the target workspace, and 'confirmed' as the staged-execution switch. It does not explain 'idempotencyKey', but that field is a standard concept and the description still covers the majority of parameters meaningfully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action — restoring a .tillbackup into a brand-new workspace — and distinguishes it from related backup tools (create_backup, verify_backup) by detailing its internal composition of create_workspace and its non-overwrite guarantee. The scope is specific and immediately understandable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit workflow guidance: the agent must first call without confirmed to receive a plan, then pass confirmed:true to execute. It also states the 'never overwrites an existing workspace' constraint, which is a clear when-not. However, it does not name alternative sibling tools (e.g., verify_backup or list_restorable_backups) directly, so it falls short of a fully explicit comparison.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resume_recurring_scheduleA
Resume a paused schedule. Periods that fell due while paused are generated by the next tick (each period separately idempotent, at most 24 per schedule per tick), which is the accepted cost D66 states. Already active settles to the same answer; ended refuses with schedule_ended.
| Name | Required | Description | Default |
|---|---|---|---|
| scheduleId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well: it discloses the catch-up behavior, the per-period idempotency, the 24-period cap per tick, and the refusal for ended schedules. It also references the accepted cost D66 states. This is meaningful beyond the name and gives an agent realistic expectations about side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and information-dense, covering purpose, edge cases, and failure modes in two sentences. It could be slightly better structured by front-loading the exact action and then the caveats, but the content is efficient and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is complete for the core action and edge cases, but lacks information about return values (no output schema), potential errors other than schedule_ended, and any permissions or preconditions. Given the relatively simple two-parameter surface, this is adequate but not comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description names neither parameter explicitly. The two parameters (workspaceId and scheduleId) are straightforward from the schema, but the description does not clarify their format, source, or constraints. Since they are simple identifiers and the schema marks them required, the description adds little beyond what any agent would assume.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a clear verb-resource pair: 'Resume a paused schedule.' This is specific enough to distinguish from siblings like pause_recurring_schedule, end_recurring_schedule, and run_due_recurring, though it does not explicitly name them. It also mentions behavioral details about catching up missed periods, which reinforces the core purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for schedules that are paused, and notes that ended schedules refuse. However, it does not explicitly state when to prefer this over similar operations like run_due_recurring or how to handle already-active schedules beyond saying they 'settle to the same answer.' There is no explicit alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retainer_burndownARead-only
Mandats-Verbrauch (B04): the drawdown/burn-down read model (P5), computed live, never a cached counter. Returns includedMinutes, carryoverInMinutes, consumedMinutes, coveredMinutes, remainingMinutes, coverageValueRappen, capRappen and overCapMinutes for a period. For a generated period it reads the frozen draw ledger; for the current in-flight period it simulates the coverage split over live approved time. An unknown retainerId returns retainer_not_found. Gated on A24 billing.read.
| Name | Required | Description | Default |
|---|---|---|---|
| periodKey | No | ||
| retainerId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds significant behavioral context beyond the readOnlyHint annotation: it clarifies that the value is computed live, never cached, and explains the simulation logic for in-flight periods. It also discloses the gating requirement (A24 billing.read) and error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise yet dense: it front-loads the core purpose, lists return fields, explains period behavior, and closes with error and gating. No wasted words, and the structure logically flows from what to how to edge cases.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description enumerates all returned fields. It covers both period scenarios, error handling, and permission gate. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has no descriptions (0% coverage), so the description compensates. It explains periodKey's role in distinguishing generated vs in-flight periods and mentions retainerId's error case. workspaceId is not explicitly explained but is self-evident as a required context parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states it is a 'drawdown/burn-down read model' and lists the exact metrics returned. It clearly identifies the resource (retainer) and the action (retrieve burndown data), distinguishing it from other retainer tools like retainer_close or retainer_generate_invoice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the behavioral difference between generated and in-flight periods, and the error case for unknown retainerId. While it doesn't explicitly name sibling alternatives, the read-only nature and metric set make it unambiguous when to use this tool versus retainer management operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retainer_closeA
Mandat beenden (B04): end a mandate (active to ended). OR 404 Abs. 1 makes a mandate terminable at any time, so close is always reachable: it refuses period_pending only while a closed period is still unbilled, and skipFinal:true overrides even that. Closing an already-ended retainer settles to the same answer (idempotent). Gated on A24 retainer.manage.
| Name | Required | Description | Default |
|---|---|---|---|
| skipFinal | No | ||
| retainerId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the required permission (A24 retainer.manage), the state transition, idempotency, and the specific refusal condition plus override. This is substantial behavioral disclosure, though it does not detail side effects beyond ending the mandate (e.g., invoicing implications).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise but dense, providing key operational details in a compact form. It front-loads the primary action and then adds conditional logic. While slightly technical, each sentence contributes useful information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, the description covers the essential aspects: when it is allowed, idempotency, permission requirements, and the main parameter nuance (skipFinal). It does not explain the return value or what 'closing' fully entails (e.g., if it triggers invoicing), but the core information needed to call it correctly is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only explains the 'skipFinal' parameter (that it overrides the period_pending refusal). It does not describe the purpose or expected values of 'workspaceId', 'retainerId', or 'idempotencyKey', which would be valuable given the lack of schema descriptions. The description only partially compensates for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb ('end') and resource ('mandate/retainer'), notes the transition from active to ended, and includes a reference code (B04). It effectively distinguishes this tool from siblings like retainer_create, retainer_update, and retainer_generate_invoice, which have different purposes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when the tool can be called: it is always reachable per OR 404 Abs. 1, with specific exceptions (period_pending while a closed period is unbilled) and how to override them (skipFinal:true). It also notes idempotency for already-ended retainers. However, it does not explicitly name alternative tools or contrast with them, though the usage context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retainer_createA
Mandat anlegen (B04): define a recurring retainer for a C00 contact (monthly|quarterly fee, included hours, optional cap, rollover). Validates feeRappen>0 (invalid_fee) and includedHours>=0 (invalid_hours) as structured refusals (P9). Returns warning:cap_below_included when the cap undercuts the included hours at the resolved rate (OP1), never an error. Mints an active retainer; posts nothing. Gated on A24 retainer.manage.
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| currency | No | ||
| rollover | No | ||
| startsOn | Yes | ||
| capRappen | No | ||
| contactId | Yes | ||
| feeRappen | Yes | ||
| projectId | No | ||
| workspaceId | Yes | ||
| includedHours | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burdenhare. It discloses validation rules (feeRappen>0, includedHours>=0), refusal format (P9), warning behavior (cap_below_included never an error), side effects (mints active retainer, posts nothing), and permission requirement (A24). This is thorough and actionable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded. The purpose appears first, followed by validation rules, warning behavior, side effects, and permission requirement. No filler or redundancy; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create tool with 6 required parameters and no output schema, the description covers the core purpose, validation, side effects, and permission. Missing details include how idempotencyKey behaves, the exact format/structure of the created retainer (e.g., return value), and semantics of startsOn or currency. These gaps are moderate but not critical for basic invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explicitly defines feeRappen and includedHours, explains the cap undercut warning, and lists period options (monthly|quarterly). It does not explain workspaceId, contactId, startsOn, idempotencyKey, currency, or projectId, but these are mostly standard identifiers whose semantics are inferable from names and required flags.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (anlegen/define) and resource (retainer) with concrete field options (monthly|quarterly fee, included hours, optional cap, rollover). Clearly distinguishes creation from other retainer tools like retainer_update or retainer_close.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool creates a new retainer, but does not explicitly name alternatives or exclusion conditions. The sibling retainer_update exists but is not referenced, so the agent must infer when to choose this tool. The A24 gating provides a precondition, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retainer_generate_invoiceA
Rechnung erzeugen (B04): turn one ENDED retainer period into an A11 invoice DRAFT (one Pauschale line plus one Zusatzaufwand line per over-cap entry), record the drawdown ledger, flip consumed B01 time to billed, and mint the rollover carryover rows, all in one transaction. Delegates to A10 createDocument (P3: no journal entry, no VAT amount, no total minted here). Idempotent per period by the (retainer, period, fee) uniqueness guard: a period already invoiced returns the existing draft with existing:true. Refuses period_not_closed, retainer_not_active, invalid_period_key, currency_mismatch before any write. Always stops at a draft (P8); gated on A24 retainer.manage.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| periodKey | Yes | ||
| retainerId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the tool is idempotent (uniqueness guard, existing:true), delegates to A10 createDocument, stops at draft (P8), and is gated on A24 retainer.manage. It also lists refusal conditions. This is detailed and goes beyond a basic 'creates an invoice' statement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured, with a clear main action, delegation note, idempotency, refusal conditions, and gating info. Each sentence adds value, and the most critical outcome (creates draft) is front-loaded. No fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex operation with 5 params, no output schema, and no annotations, the description covers the input requirements (via refusal conditions), the transaction scope, the side effects, and the idempotency behavior. It also mentions performance gating. Missing parameter-level syntax is minor given the rich procedural detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for all 5 parameters, so the description must compensate. It mentions periodKey, idempotencyKey, and retainerId implicitly through the flow, but doesn't explain the format or purpose of each parameter explicitly. The description references the uniqueness guard involving retainer, period, fee, which gives some context for periodKey and retainerId, but actor and workspaceId are not explained. Baseline 3 is fair given the partial coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it converts an ended retainer period into an invoice draft, lists specific line items (Pauschale and Zusatzaufwand), and mentions the ledger, billing, and rollover effects. It distinguishes from related operations like billing_generate_invoice and retainer_close by focusing on the retainer-specific invoice generation flow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states the conditions for use (period must be ended) and refusal reasons (period_not_closed, retainer_not_active, etc.), which guides when to call. However, it doesn't explicitly name alternatives or say when NOT to use this tool vs others like billing_generate_invoice, but the context of retainer period invoicing is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retainer_listARead-only
Mandate anzeigen (B04): the Mandate list (P5), tenant-scoped, filterable by contactId and status, with the G00 saved-view seam (OP10, savedViewId). Gated on A24 billing.read.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| contactId | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true. The description adds the billing.read permission gate and tenant scoping, which are useful behavioral details beyond the structured annotations. It does not discuss pagination or return shape, but the annotation lowers the bar for that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core purpose. Internal codes like B04, P5, G00, OP10, and A24 add noise and are unexplained, which slightly reduces clarity, but every sentence still contributes meaningful context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with no output schema and no param descriptions, the description provides enough to invoke the tool correctly. However, it does not cover return shape, pagination, or how savedViewId interacts with contactId/status filters, leaving some ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of parameter meaning. It explains contactId and status as filters, and savedViewId as the saved-view seam. The required workspaceId is implied by 'tenant-scoped', though not explicitly named.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('list' / 'Mandate anzeigen') and resource ('Mandate list'), and adds tenant scope and filter parameters (contactId, status, savedViewId). This clearly distinguishes it from sibling operations like retainer_create, retainer_update, and retainer_close.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It conveys that the tool is tenant-scoped and filterable by contactId/status, which implies when it should be used. However, it does not explicitly mention alternatives or exclusions, such as when to prefer a sibling tool like retainer_burndown or retainer_generate_invoice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retainer_run_dueA
Alle fälligen abrechnen (B04): bill every active retainer whose period ended before asOf, iterating the retainer table directly (never an A12 recurring_schedule). SELF-KEYED (no idempotencyKey): idempotent per period via the (retainer, period, fee) uniqueness guard, so re-running the tick at any cadence bills each period exactly once; each period is its own transaction. Agent/scheduler/bulk-action oriented. Gated on A24 retainer.manage; denylisted from automation (a bulk generation tick a human or cron owns).
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| actor | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full responsibility and delivers richly. It discloses idempotency behavior ('SELF-KEYED (no idempotencyKey)' with a uniqueness guard), transaction granularity ('each period is its own transaction'), and the direct iteration source. It also states permission requirements ('Gated on A24 retainer.manage') and automation restrictions, giving an agent a clear behavioral model.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-organized, with no filler words. It front-loads the primary action and then adds key behavioral constraints in a structured way. Slightly overloaded with internal codes (B04, A12, A24), but these appear to be meaningful domain references, not noise. Every sentence contributes useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a bulk billing tick, the description covers the main concerns: triggering condition, idempotency, transaction isolation, permission, and automation restrictions. It doesn't describe return values (no output schema exists), but that's likely not critical for a side-effect oriented operation. Missing details like the exact format of 'asOf' or behavior when no periods are due are minor given the robust behavioral disclosure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all parameters. It only clarifies 'asOf' ('period ended before asOf'), but says nothing about 'workspaceId' or 'actor'. While workspaceId is common, this is a critical gap for a tool with no schema-level documentation. The description does not explain the actor parameter or its role, leaving two of three parameters underspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb ('bill'), a specific resource ('every active retainer whose period ended before asOf'), and a clear scope ('iterating the retainer table directly'). It explicitly differentiates itself from A12 recurring_schedule, making its purpose unambiguous and distinguishable from sibling tools like retainer_generate_invoice and recurring_schedule tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides direct usage context: it is 'Agent/scheduler/bulk-action oriented', 'denylisted from automation', and 'a bulk generation tick a human or cron owns'. It also clarifies the exclusion: 'never an A12 recurring_schedule'. It does not name a specific alternative tool, but the differentiation and operator guidance are sufficient for an agent to decide when to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retainer_updateA
Mandat bearbeiten (B04): patch a retainer. Coverage terms (feeRappen, includedHours, capRappen, rollover) apply to FUTURE periods only, which is structural: a generated period stored its own draws and generation reads the row live, so an edit never rewrites history. The identity fields (period, contactId, projectId) freeze once any draw exists (retainer_has_draws). Gated on A24 retainer.manage.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | No | ||
| retainerId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden and delivers: it discloses that coverage terms apply to future periods only, that an edit never rewrites history (with the structural reason), and that identity fields freeze once any draw exists. This is exactly the behavioral context an agent needs for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the purpose, then dense behavioral semantics. The structural explanation is verbose but earns its place because it justifies a non-obvious invariance. No filler or repetition of schema trivia.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a moderately complex mutating tool with no output schema and no annotations, the description covers the critical invocation context: effect scope, identity freeze, and permission gating. The one notable gap is the required idempotencyKey — the agent is told to send it but not what idempotency guarantees it provides.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the patch parameter is an opaque, untyped object, so the description must compensate. It does by enumerating the patchable fields (feeRappen, includedHours, capRappen, rollover) and clarifying the identity fields (period, contactId, projectId). It stops short of describing patch structure, types, or whether all fields are optional, so it is not a full 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
"Mandat bearbeiten (B04): patch a retainer" states a specific verb (patch) and resource (retainer), and the retainer_* sibling family (create/close/list/run_due) makes the update role unambiguous. The description also names the exact fields affected, so an agent knows precisely what this tool operates on.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than explicit: as an update tool it presumably applies to existing retainers, and the permission gate A24 retainer.manage is stated. However, no sibling is named and there is no explicit when-to-use vs retainer_create/retainer_close guidance or exclusion conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
retry_automation_runA
Finish a run that is stuck in running, which is what a process death between the claim and the settle leaves behind. Re-sends the SAME stored input, including the derived idempotency key, so a lost invocation that had in fact committed cannot happen twice. Runs as the actor the firing ran as, so a since-demoted author is refused now exactly as it would be on a fresh firing. Requires manage_automations. A settled run refuses with run_not_stuck.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses idempotency behavior (re-sends same input with derived idempotency key to prevent double-commit), actor context (runs as the firing actor), and the error case for settled runs. It also states the permission requirement. This is exceptional transparency for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph but dense with information: purpose, cause, idempotency behavior, actor context, permission, and refusal condition. It is well-structured and front-loads the main action, though it could be slightly more concise. It earns its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core purpose, behavioral guarantees, error conditions, and permissions. It does not describe the return value (no output schema), but for a retry operation this is often not critical. Given the tool's complexity and the absence of annotations, it is quite complete. A small gap is not explaining what happens on success or how the result is returned.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions the idempotency key as derived and stored, giving meaning beyond the schema. workspaceId and runId are self-evident from the name and context, but the description does not explicitly describe them. However, it provides enough context for an agent to infer their purpose. A score of 4 is appropriate given the 0% coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Finish a run that is stuck in `running`' – a specific verb and resource. It also distinguishes from sibling tools by explaining the exact condition it addresses and what it does differently (re-sends the same stored input). This is precise and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use it (when a run is stuck in 'running' due to a process death) and when it will refuse (settled run with run_not_stuck). It also mentions the permission requirement (manage_automations) and the actor context, which helps the agent decide if it's appropriate. However, it does not explicitly name alternative tools for other scenarios, but the context is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reverse_entryA
Reverse a posted entry by posting its faithful mirror (OR 957a). Refuses owned_by on an entry a capability minted with a row of its own (an MWST-Saldierung, an Abgrenzung, a Rückstellung and their mirrors): reverse those through the owning verb (vat_settlement_reverse, accrual_reverse, provision_reverse, provision_release_reverse) so the row moves with the entry. CONSEQUENCE: Posts an immutable reversing entry against a posted one; it cannot be un-posted.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | ||
| entryId | Yes | ||
| description | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully carries the burden of behavioral disclosure. It explicitly states the consequence: 'Posts an immutable reversing entry against a posted one; it cannot be un-posted.' This is critical behavior for an agent to know. It also discloses the refusal condition and the logic behind it, going beyond a simple action description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and technical, but every sentence serves a purpose: it states the action, the refusal cases with alternatives, and the irreversible consequence. It is front-loaded with the main purpose and packs necessary exceptions. While it could be more concise, it avoids fluff and is efficiently structured for an agent that needs precise domain-specific details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex accounting tool with no annotations or output schema, the description covers the essential behavioral context: what it does, what it refuses, and its irreversible nature. It does not explain parameter semantics (which is a gap), but given the tool's specificity, it provides sufficient context for an agent to understand the primary use case and constraints. The missing parameter details are a notable omission but not fatal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 5 parameters with 0% description coverage, yet the tool description does not explain any of them. While it implicitly refers to 'entryId' as the entry to reverse, it does not clarify the roles of workspaceId, idempotencyKey, date, or description. For an irreversible posting action, idempotencyKey is especially important and should be explained. The description adds minimal meaning to the parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Reverse a posted entry by posting its faithful mirror') and the resource (a posted entry). It also distinguishes itself from sibling reversal tools by explicitly naming the cases it refuses and the alternative verbs to use. This makes the tool's purpose unambiguous and distinct from related tools like vat_settlement_reverse, accrual_reverse, etc.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when NOT to use this tool: it refuses entries that have a 'row of its own' (MWST-Saldierung, Abgrenzung, Rückstellung and their mirrors) and directs the agent to use the owning verb instead. This gives clear guidance on when to choose this tool versus alternatives, which is exactly what usage guidelines should provide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reverse_paymentA
Reverse a payment (storniert eine Zahlung): posts a reversing entry and re-opens every document it settled. It never deletes, the original entry stays untouched, and a payment is always reversed as a whole. Requires intent='reverse_payment'. CONSEQUENCE: Posts a reversing entry against the payment and reopens what it had settled.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | ||
| intent | Yes | ||
| paymentId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does this well: it explains the mechanism (posts a reversing entry), the side effect (re-opens every document it settled), the non-destructive nature (never deletes, original entry stays untouched), the all-or-nothing scope (always reversed as a whole), and the required intent parameter. This is exemplary transparency for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the main action and consequence appear in the first sentence, followed by key clarifications (non-deletion, whole-payment scope, intent requirement). Every sentence adds value; there is no filler or repetition of schema field names.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers the essential behavioral contract: what happens, what doesn't happen, and the required intent. It does not describe the return value or error conditions, but those are not strictly necessary for an agent to invoke the tool correctly. The main gap is not explaining the idempotencyKey semantics, but the parameter name is self-explanatory.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explicitly mentions the 'intent' parameter and its required value ('intent='reverse_payment''), which is critical. It does not explain the other parameters (workspaceId, paymentId, idempotencyKey, date), but their names are self-explanatory and the description's mention of 'payment' and 'reversing entry' gives enough context. The intent requirement is the most important semantic detail and it is covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Reverse'), a specific resource ('a payment'), and the core behavior: posts a reversing entry and re-opens every document it settled. It also explicitly distinguishes itself from deletion ('It never deletes, the original entry stays untouched'), which helps differentiate it from sibling tools like reverse_entry or void_vendor_bill.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when to use this tool: to reverse a payment as a whole, with the consequence of reopening settled documents. It does not explicitly name alternative tools or state when not to use it, but the context is clear enough for an agent to select it over related tools like reverse_entry or void_vendor_bill.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_bank_txnA
Signal that ONE booked bank debit needs a human's review because the ranked suggestion found no candidate at or above the workspace threshold (A36). Writes no ledger row and books nothing: it is a pure signal whose success emits bank_txn.needs_review for a G01 rule to react to (e.g. create an E03 task). Idempotent per (bankTxnId, idempotencyKey). A credit-classified txn refuses with use_qr_queue (the credit review moment is A21's qr_match.needs_review). A caller runs the suggestion pass and calls this once per booked debit, so N unmatched txns fire N occurrences.
| Name | Required | Description | Default |
|---|---|---|---|
| bankTxnId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses behavior: writes nothing, is a pure signal, emits bank_txn.needs_review, is idempotent, refuses credits with use_qr_queue, and scales linearly with unmatched bookings. This covers side effects, error behavior, and operational semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place, with key information front-loaded: purpose, non-side-effect, idempotency, credit refusal, and usage pattern. No redundancy or filler; compact but information-dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with no output schema and no annotations, the description covers all critical aspects: when to call, side effects, idempotency, error mode, trigger condition, and integration with rules (A36/G01/A21). An agent gains sufficient knowledge to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains the roles of bankTxnId and idempotencyKey (per (bankTxnId, idempotencyKey)) and implies workspaceId is the workspace context. Does not detail types/formats, but provides enough semantic meaning beyond the bare schema for the core parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (signal review) on a specific resource (booked bank debit) with a precise trigger condition (no candidate at/above workspace threshold, A36). Distinguishes from credit matching by naming the alternative mechanism (A21's qr_match.needs_review) and clarifies it is not a write/booking operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes when to call (after suggestion pass, once per booked debit, for N unmatched txns) and when not to call (credit-classified txns refuse with use_qr_queue). Also states idempotency per (bankTxnId, idempotencyKey) and the expected frequency (N occurrences).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_statusARead-only
Review coverage for a period ('YYYY-MM' or 'YYYY'): approved/flagged/open counts plus every posted entry with its current review state, last reviewer and comment/flag counts. This is the bar behind the review surface, and what says whether a period is ready for lock_period (A03) and export. savedViewId applies a saved review view (G00): its stored period is used when none is named here; an explicit period always wins.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark readOnlyHint=true, and the description is consistent with that. It adds useful behavioral detail beyond the annotation: period format support, savedViewId precedence, and the fact that an explicit period wins over a saved view's period. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core return values, and every sentence earns its place: output semantics, integration context, and parameter behavior. No filler or redundant restatement of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, the description appropriately discloses return contents and connects the tool to related workflows (lock_period/export). It does not mention pagination or filtering, but for this read-only review-coverage tool the provided information is sufficient for an agent to select and invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden. It compensates well for period and savedViewId by explaining the accepted formats ('YYYY-MM' or 'YYYY') and the precedence rule. workspaceId is left implicit, but it is a common scoping parameter and is marked required in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific verb-resource pair: 'Review coverage for a period,' and enumerates exactly what is returned (approved/flagged/open counts, posted entries, review state, last reviewer, comment/flag counts). It gives enough identity to distinguish it from most siblings, though it does not explicitly name an alternative tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear usage context: this tool is 'the bar behind the review surface' and determines whether a period is ready for lock_period (A03) and export. It stops short of listing when-not-to-use conditions or naming alternative review-related tools, but the intended trigger is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
revoke_memberA
Remove a member's access to this workspace. The person's identity survives, so re-inviting them restores their history. Revoking the only remaining owner is refused (last_owner). CONSEQUENCE: Removes a member's access to the workspace; the lockout applies on their next call.
| Name | Required | Description | Default |
|---|---|---|---|
| memberId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It covers identity retention, re-invite history restoration, the last-owner refusal, and lockout timing on next call. It does not cover permission requirements or audit-log effects, but the disclosed behaviors are meaningful and beyond what structured fields would show.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, with each sentence adding distinct value: the action, the re-invitation consequence, the last-owner refusal, and the lockout timing. The scannable CONSEQUENCE label and tight prose make it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation tool without an output schema, the description covers the action, side effects, edge-case refusal, and timing. Missing details like whether tokens or API keys are revoked would be ideal, but the provided information is sufficient for correct invocation in most scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. It implies that workspaceId identifies this workspace and memberId identifies the member, but it does not explain their format, source, or validation requirements, leaving both required parameters under-documented at both schema and description level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Remove'), a resource ('member's access to this workspace'), and adds behavioral scope (refuses last-owner revocation). This distinguishes it from sibling tools like archive_account or portal_grant_revoke.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool (when a member's access must be removed) and mentions the re-invitation alternative for restoring history, providing useful context. However, it does not explicitly reference alternative tools or state when not to use it in favor of role changes or workspace archival.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_due_automationsA
The tick for schedule triggers: fire every enabled rule whose cadence has come due at asOf, at most ONCE each however far behind it had fallen. Requires manage_automations, because the tick is what makes a schedule rule write unattended. asOf may name a past instant but never a future one: the injected clock decides what is due, not the caller. Safe to call as often as you like: a repeated tick for the same asOf computes the same occurrence key and does nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers substantial behavior: the tick causes unattended writes, requires manage_automations, uses an injected clock rather than caller control, and is idempotent via the same occurrence key. It does not disclose what the call returns or what happens if an individual rule fails, so it is not fully complete, but it is far beyond minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, each earning its place. The core action and scope are front-loaded, and the permission requirement, clock semantics, and idempotency follow without repetition or padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a side-effecting automation tick with no annotations and no output schema, the description covers purpose, permission, temporal semantics, and idempotency. The main gap is that it does not explain the return value or failure behavior when rules run, which an agent might need to interpret the result. Still, it is quite complete for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It thoroughly explains asOf: it can name a past instant, never a future one, and the injected clock determines due-ness. workspaceId is not described, but its role as a scoping identifier is reasonably self-evident. Overall the description adds real semantic meaning beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the exact operation: 'fire every enabled rule whose cadence has come due at asOf'. It states the resource (schedule triggers/automation rules), the action (tick/fire), and the key constraint (at most once each even if far behind). This clearly differentiates it from sibling tools like retry_automation_run or list_automation_runs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when to call: when a schedule rule's cadence is due. It also states a prerequisite (manage_automations permission), temporal constraints (asOf may be past but never future), and that it is safe to call repeatedly. It does not explicitly name alternatives or say when not to use it, but the guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_due_recurringA
The tick for Serienrechnungen: settle every due period of every active schedule at asOf (default now; a future asOf is refused; a past one only bounds how far the catch-up reaches, it never dates anything). Each occurrence INVOKES the registered create_document, and issue_invoice when the schedule has autoIssue, through the shared dispatch as the schedule AUTHOR, with a key derived from (scheduleId, periodKey), so a re-tick never double-bills (§H-IDEMPOTENT on rows). Every generated line carries its period as the Leistungsdatum, so a catch-up bills each period at its own VAT era; the due date is the clock day plus dueDays. A locked target period keeps the draft, records skipped_locked and is retried after unlock (§H-PERIOD); a waiting draft a human has since ISSUED settles the period as issued, and one a human has since CANCELLED settles it as discarded, so both operator responses converge and neither can strand a schedule or its siblings. Catch-up generates each missed period, at most 24 per schedule per tick.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so excellently. It discloses side effects (invoking create_document and issue_invoice), idempotency (re-tick never double-bills), edge cases (locked periods, waiting drafts issued or cancelled), asOf semantics (future refused, past bounds catch-up), and the catch-up limit. It even explains how operator responses converge for locked/cancelled drafts. This is exceptionally transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence adds unique value. It is front-loaded with the core purpose, then covers asOf behavior, invocation of downstream tools, idempotency, edge-case handling, and catch-up limits. There is no fluff; it is a model of concise yet comprehensive documentation for a complex operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (side effects, idempotency, multiple edge cases, catch-up logic), the description is remarkably complete. It covers all behavioral aspects an agent needs to know before invoking, including what happens with locked periods, cancelled drafts, and the catch-up limit. There is no output schema, but the description does not need to describe return values; it focuses on the operation's semantics, which is sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has only workspaceId (required) and asOf (optional), with 0% description coverage. The description compensates for asOf by explaining its semantics in detail (default now, future refused, past only bounds catch-up). For workspaceId, it does not explain it, but that is a standard workspace-scoping parameter across many tools and is likely self-evident. Still, a brief mention of workspaceId's role would have been beneficial, but the description covers the most critical parameter thoroughly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's verb and resource: 'settle every due period of every active schedule'. It precisely defines the scope (all active schedules) and the action (settling due periods). It also mentions the default asOf and the behavior for future/past asOf, which further distinguishes it. Among siblings, it stands out as the runner for recurring schedules, not a CRUD operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly conveys when to use it: when you need to run due recurring invoices. It explains the catch-up behavior, the limit of 24 periods, and the asOf constraints. However, it does not explicitly contrast with sibling tools like run_due_automations or create_recurring_schedule, but the purpose is clear enough that an agent would know to invoke this when it's time to process recurring schedules. A brief note on when not to use it (e.g., for single manual invoice generation) would be a minor improvement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
runtime_catalogBRead-only
Die angebotenen Modelle (P5): the static manifest the installed companion package shipped, a pure read of a local file and never a fetch (new models arrive via npm update). Per row the plain-language German-quality sentence, RAM floor, download size and licence; rows over the RAM of this machine come back fits:false with the have/need figures, and exactly one fitting row is recommended and preselected.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses that the tool reads a local manifest rather than fetching, that rows over the machine's RAM return fits:false with have/need figures, and that exactly one fitting row is recommended and preselected. This adds meaningful behavioral detail beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense, run-on, mixed German/English sentence with parenthetical asides and semicolons. It contains useful information but is poorly structured and not front-loaded for quick agent comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers output-row semantics reasonably well but leaves the sole required parameter entirely unexplained. With no output schema and no parameter documentation, the agent cannot confidently construct a correct call without additional inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one required parameter, workspaceId, with 0% schema description coverage, and the description never explains what workspaceId means, its format, or how it affects the result. Since schema coverage is low, the description had the burden to compensate but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool exposes the offered models (P5) from a static manifest, with per-row details like RAM floor, download size, licence, and fit status. This conveys a catalog/listing purpose and distinguishes it from network-fetching behavior, though it never explicitly names a sibling like runtime_select to sharpen the contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear contextual signals: the result is a pure read of a local file, never a fetch, and new models arrive via npm update. However, it does not explicitly say when to prefer this over runtime_status/runtime_select, nor does it state any prerequisites beyond the required workspaceId.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
runtime_selectA
Wähle das Modell der lokalen Entwurfs-Engine: persists the choice of the workspace from the shipped catalog (source=catalog, validated against the manifest AND the RAM floor of the machine IN THE VERB, so an agent cannot select a model the machine cannot run: unknown_model_ref retains the old selection, insufficient_ram names the have/need GB) or a local .gguf path behind Erweitert (source=byo, unsupported, no quality claim). Reversible and visible in runtime_status.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes | ||
| ggufPath | No | ||
| modelRef | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers: validation against manifest and RAM floor, precise error outcomes (unknown_model_ref retains old selection, insufficient_ram reports have/need GB), unsupported byo path with no quality claim, and reversibility. This is exemplary behavioral disclosure beyond a generic 'select model' statement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense run-on sentence mixing German and English, with nested parentheticals such as 'validated against the manifest AND the RAM floor of the machine IN THE VERB' and 'unknown_model_ref retains the old selection, insufficient_ram names the have/need GB'. While information-rich, it is poorly structured and hard to parse; a cleaner multi-sentence format would improve usability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations or output schema, the description provides strong coverage of purpose, sources, validation behavior, error states, and reversibility. It lacks explicit mapping of all parameters (notably idempotencyKey) and return format, but the behavioral detail compensates for most of what the structured fields do not provide.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does clarify source=catalog vs source=byo, thereby giving meaning to modelRef and ggufPath. However, idempotencyKey and workspaceId are not described; workspaceId is inferable from its name, but idempotencyKey is left entirely to the parameter name with no guidance on its role.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states it selects and persists the model choice for the workspace, with two modes: source=catalog and source=byo. The verb 'persists' plus the resource 'model of the local Entwurfs-Engine' are specific, but the description does not explicitly distinguish this tool from runtime_catalog or runtime_status, though it references the latter for reversibility.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives internal usage guidance for choosing between catalog and byo sources, including validation and unsupported status, but it does not explicitly say when to use this tool versus alternatives like runtime_catalog or runtime_status. Cross-tool routing is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
runtime_statusARead-only
Der Status der lokalen Entwurfs-Engine (P5): whether an OP6 adapter is registered, which runtime, model and device, why a load failed if it did, and the persisted model selection of the workspace. With nothing installed the answer is registered:false, and no cloud path exists to fall back to.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond the readOnlyHint annotation by disclosing the exact output aspects, including failure reasons and the behavior when nothing is installed (registered:false with no cloud fallback). This is valuable behavioral context, though it could further describe error cases or response shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences with the main purpose front-loaded and specific technical details following. Every clause adds information about the returned status or edge-case behavior, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main return dimensions and a key edge case, which is substantial for a single-parameter read-only status tool with no output schema. It does not mention all possible failure modes or the meaning of the acronyms, leaving some room for improvement.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description must compensate for the single workspaceId parameter. It does tie the parameter to the workspace by mentioning the 'persisted model selection of the workspace', so the agent understands the parameter's role, but it does not specify the format, source, or any constraints around workspaceId.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the resource as the local design engine (P5) and enumerates exactly what status information is returned: adapter registration, runtime, model, device, load failure reasons, and persisted model selection. The meaning of acronyms like P5 and OP6 is not explained, and no sibling differentiation is stated, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It is clearly a status/diagnostic tool, and the note about 'nothing installed' giving 'registered:false' and having no cloud fallback provides useful context about when to consult it. However, it does not explicitly mention alternatives such as runtime_catalog or runtime_select, nor gives when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sales_order_backordersARead-only
Liste die Lieferrückstände (P5): every order line with backorder_qty > 0, joined against current D01 on-hand so you see which backorders are coverable now. An agent can poll this after a D02 receipt to trigger a second delivery. An empty list is not an error.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only safety profile (readOnlyHint=true), so the bar is lower. The description adds genuine behavioral value beyond the annotation: the join semantics showing which backorders are coverable now, the polling pattern, and the explicit guard that an empty list is not an error – which prevents an agent from misreporting a normal result as failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the definition, the usage trigger, and the empty-result guard. No filler. The German opening followed by English text is a minor structural inconsistency that keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with a single scoping parameter, no output schema, and a readOnlyHint annotation, the description covers query semantics, the triggering event, and the empty-result behavior. It omits the return shape and pagination, but with no output schema and a self-evident list return those are minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description never mentions workspaceId, so it fails to compensate for the structured-data gap. The parameter is trivial – a standard workspace scoping string self-evident from its name and the sibling set – but the description itself contributes nothing about it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource statement – 'Liste die Lieferrückstände (P5)' – and sharpens it with the exact filter (backorder_qty > 0) and the D01 on-hand join, so an agent knows precisely what the tool returns. This clearly distinguishes it from siblings like po_open_lines, stock_on_hand, and sales_order_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger: poll this after a D02 receipt to trigger a second delivery. That is clear when-to-use context. It stops short of a 5 because it names no alternatives or when-not-to-use exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sales_order_cancelA
Storniere einen Auftrag (draft|confirmed -> cancelled): keeps open-order lists honest. An order with any ISSUED delivery note is refused with has_deliveries, because shipped goods come back only through an explicit D01 return movement, never by erasing the order (H-AUDIT).
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| salesOrderId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses the key behavior: the cancellation, the refusal when deliveries exist, and the audit rationale (H-AUDIT) that shipped goods must be returned via an explicit movement. It doesn't describe idempotencyKey behavior or success response, but the main side effects and constraints are well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the primary action and state transition, then adding the refusal condition. Every word earns its place; there is no redundancy or filler. The structure is efficient and immediately understandable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is complete for the core action: it tells the agent what it does, the allowed states, and the key refusal condition. It lacks details on idempotencyKey usage and return values, but given the simplicity of the operation and the absence of an output schema, the agent can likely proceed. The missing idempotency guidance is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does not explain any of the three parameters (workspaceId, salesOrderId, idempotencyKey). While the names are self-explanatory, the description provides no additional meaning, such as the role of idempotencyKey or any parameter-specific constraints. This is a significant gap given the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Storniere' (cancel) and the resource 'Auftrag' (order), and specifies the exact state transition (draft|confirmed -> cancelled). This distinguishes it from sibling tools like sales_order_confirm or sales_order_invoice, which operate on different states. The purpose is unambiguous and specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear conditions for use: cancel an order in draft or confirmed state, and it explicitly warns that orders with issued delivery notes are refused (has_deliveries). This effectively tells the agent when NOT to use it (if deliveries exist) and points to the alternative of a D01 return movement instead of cancellation. It doesn't name specific sibling tools but provides adequate usage guidance through the state and refusal conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sales_order_confirmA
Bestätige einen Auftrag (draft -> confirmed): snapshots stock allocation. Every non-stock (service, free-text) line is auto-delivered on confirmation with NO stock movement, which is what makes it invoiceable without a delivery note; a stock line gets backorder_qty = max(0, ordered - D01 on-hand). A pure-service order advances straight to delivered, a mixed order to partially_delivered, an all-stock order stays confirmed. An empty order is refused (no_lines); a non-draft order is refused (invalid_transition).
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| salesOrderId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is exceptionally transparent, describing exact stock allocation behavior (backorder_qty formula), auto-delivery of non-stock lines with no stock movement, resulting order statuses for different compositions (pure-service, mixed, all-stock), and refusal conditions (no_lines, invalid_transition). Since no annotations are provided, this rich detail fully carries the burden of behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph but each clause carries meaningful information about behavior, transitions, or errors. It is front-loaded with the core purpose and then details edge cases, making it informative without being overly verbose for the complexity involved.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the operation, the description covers all major outcomes: status changes, stock handling, and refusal conditions. It does not explain the meaning of 'D01 on-hand' or describe the response format, but no output schema exists and the primary invocation parameters are clear. The coverage is strong overall.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, so the description was expected to compensate by explaining parameterscars. However, it only mentions the order (presumably salesOrderId) but does not clarify workspaceId or idempotencyKey. The parameter names are somewhat self-explanatory, but no additional semantic value is added beyond the schema's property names and types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action as confirming a sales order with the transition 'draft -> confirmed' and adds domain-specific behavior such as snapshotting stock allocation. It distinguishes itself from siblings like sales_order_create, sales_order_cancel, and sales_order_invoice by detailing the confirmation semantics and stock/status outcomes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: on a draft sales order, and explicitly notes that empty or non-draft orders are refused with specific error codes. However, it does not explicitly name alternatives or state 'use this instead of X', though the context makes the usage clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sales_order_createA
Lege einen Auftrag an: a sales order on a C00 contact with D00 item lines. Each item line is priced once through D00 (contact list, segment list, item base) and its tax code resolved once through the item A05 default, then both are snapshotted as literals on the so_line (the freeze, H-VAT-TRACE). POSTS NOTHING and touches no stock: a sales order is a confirmed demand, not a document. A EUR order snapshots the CHF/txn fx_rate for the H-FX trace. An order with zero lines is a legal draft (confirm refuses an empty order); a foreign contact/item id is refused (invalid_reference / unknown_item).
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| lines | No | ||
| notes | No | ||
| quoteId | No | ||
| currency | No | ||
| contactId | No | ||
| expectedOn | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility for behavioral disclosure. It thoroughly discloses side effects: 'POSTS NOTHING and touches no stock', it explains the pricing and tax freezing mechanism ('snapshotted as literals'), describes the fx handling for EUR orders, and specifies edge cases (zero-line draft, foreign id rejection). This goes well beyond a basic 'creates a sales order' and gives the agent a precise mental model of what happens.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph but every sentence adds value. It front-loads the primary action and then covers pricing, side effects, fx, and validation rules. While not broken into structured sections, it is concise and information-dense without redundancy. The only minor issue is that it is a wall of text that could benefit from bullet points, but it remains efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is comprehensive regarding behavior and validation rules, but it omits guidance on how to use the parameters. For a tool with 9 parameters, the absence of any mapping from description to schema leaves a significant gap. The description does not explain, for example, the role of 'quoteId' or 'idempotencyKey'. It covers the 'what' and 'why' but not the 'how' of parameter usage, making it incomplete for a complex creation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, meaning the description does not explain any of the 9 parameters (actor, lines, notes, quoteId, currency, contactId, expectedOn, workspaceId, idempotencyKey). While the description mentions concepts like 'contact', 'item lines', and 'tax code', it never maps these to the schema parameter names, so an agent cannot infer what each field expects. The description adds domain context but fails to compensate for the complete lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Lege einen Auftrag an: a sales order on a C00 contact with D00 item lines', which clearly states the action (create), the resource (sales order), and the context (C00 contact, D00 item lines). It differentiates itself from siblings like sales_order_confirm and sales_order_cancel by explicitly noting this creates an order that is later confirmed, and it distinguishes from posting tools by stating 'POSTS NOTHING'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides context about when this tool is appropriate: it creates a draft that can be confirmed (implying a workflow with sales_order_confirm), and it notes that a zero-line order is a legal draft, which hints at validation. However, it does not explicitly name sibling tools like sales_order_from_quote or sales_order_invoice as alternatives, nor does it state conditions for when not to use this tool. The guidance is implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sales_order_from_quoteA
Wandle eine angenommene Offerte in einen Auftrag um: reads the accepted C02 quote and copies its FROZEN lines (item, qty, price, tax_code travel unchanged, the VAT trace) into a new draft order linked by quote_id. Idempotent and single-order: converting the same quote twice returns the FIRST order, never a second. A non-accepted quote is refused (quote_not_accepted).
| Name | Required | Description | Default |
|---|---|---|---|
| quoteId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description carries the full burden and does a good job disclosing idempotent and single-order behavior ('converting the same quote twice returns the FIRST order, never a second') and the refusal behavior. It also notes that lines are frozen lines and VAT trace travel unchanged, which is genuinely useful beyond what annotations would have provided (none exist). The only gap is that side effects on the source quote are not described; the description says 'draft order' but doesn't state if the accepted quote is marked as converted or whether anything else changes. Minor side-effect disclosure missing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and information-dense, front-loaded with the operation, then idempotency, then rejection behavior. It loses a point because the opening clause 'Wandle eine angenommene Offerte in einen Auftrag um' is a German restatement of the tool name, which adds little beyond the English text that follows.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description does not describe the return value (the order ID), which an agent must know to verify the result. The param semantics are incomplete (workspaceId is unexplained, idempotencyKey only implied. For a tool with moderate complexity, the missing return value and workspaceId make the description not fully complete for invocation and post-call verification.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for all three parameters. It only touches quote_id ('linked by quote_id'), implicitly connecting to the quoteId parameter, and the idempotencyKey is only loosely implied by the idempotency statement. workspaceId is completely unexplained. With 0% schema coverage and 3 parameters, this is a clear gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific, well-scoped operation: converting an accepted quote into a draft order, and specifies exactly what is copied (item, qty, price, tax_code, VAT trace), which distinguishes it from siblings like sales_order_create or quotes_convert. The description adds a precise pre-condition (accepted quote) and distinguishes from a generic create tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context: use only when the quote is accepted, and non-accepted quotes are refused with quote_not_accepted. It doesn't name alternatives explicitly, but the precondition is strong enough to route an agent to this tool for accepted-quote-to-order conversion without opening siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sales_order_getARead-only
Lies einen Auftrag mit Positionen (ordered / delivered / invoiced / backorder qty), seinen Lieferscheinen und den verknüpften Rechnungen. The one read that shows the whole order -> delivery -> invoice chain for a single order. savedViewId is the G00 coverage seam over the delivery_note kind (its custom-field columns surface in the order detail).
| Name | Required | Description | Default |
|---|---|---|---|
| savedViewId | No | ||
| workspaceId | Yes | ||
| salesOrderId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true and the description is consistent with it ('Lies'/'read'). Beyond the annotation, it discloses what the tool aggregates — positions with per-status quantities, delivery notes, and invoices — and explains how savedViewId surfaces delivery_note custom-field columns in the order detail. This adds meaningful context without contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences each do real work: what is returned, how the tool is positioned among siblings, and what the optional parameter does. The mix of German and English is slightly awkward and 'G00 coverage seam' is jargon, but there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-order read with no output schema, the description inventories the return payload (positions with qty breakdown, delivery notes, linked invoices) and resolves the savedViewId nuance. It falls short of a 5 because it never describes the response shape or how the delivery/invoice records are nested.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the full burden. It genuinely explains savedViewId as a 'G00 coverage seam over the delivery_note kind' whose custom-field columns surface in the order detail, but workspaceId and salesOrderId, though likely inferable from names and the required list, get no descriptive support.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Lies einen Auftrag') and enumerates exactly what is returned: positions with ordered/delivered/invoiced/backorder quantities, delivery notes, and linked invoices. It then distinguishes itself from the sales_order_* siblings by positioning itself as 'The one read that shows the whole order -> delivery -> invoice chain for a single order,' making it easy to tell apart from sales_order_list and the mutation tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'The one read that shows the whole order -> delivery -> invoice chain for a single order' gives clear context for when to select this tool over a list or status read. It does not explicitly name an alternative or state when not to use it, so it stops short of full exclusion guidance that would earn a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sales_order_invoiceA
Erstelle eine Rechnung aus einem Auftrag (P8, lands as an A11 DRAFT): collects the delivered-but-uninvoiced qty (including confirm-time auto-delivered service lines) and delegates to A10 createDocument (type invoice). D03 mints NO journal entry and stores no totals: A11 -> A02 own the posting at issue (P3). Each invoiced portion is one so_line_invoice link row, so a line invoiced across two partial invoices carries two rows; the invoiceable remainder is pre-checked, so a fully-invoiced line writes ZERO rows and returns nothing_to_invoice. NO DOUBLE-BILLING.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| workspaceId | Yes | ||
| salesOrderId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations to lean on, the description carries the full burden and does so thoroughly. It discloses draft status, absence of journal entry/postings, per-line link row behavior, partial-invoice row duplication, the nothing_to_invoice zero-row result, and NO DOUBLE-BILLING.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded, with every sentence contributing substantive behavioral detail. However, the heavy use of unexplained domain codes (P8, A11, D03, A10, A02, P3) and mixed German/English phrasing reduces readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the workflow, draft lifecycle, partial-invoicing behavior, and an empty-result edge case, which is strong for a no-output-schema tool. It still omits the success return value shape, explicit parameter semantics, and any alternative-tool routing guidance such as when issue_invoice would be the correct sibling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description only implicitly maps 'salesOrderId' via 'Auftrag and never explains the remaining parameters. actor and idempotencyKey are left entirely to inference, and no parameter-specific guidance is provided despite the full coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Erstelle eine Rechnung aus einem Auftrag') and clarifies it lands as an A11 DRAFT. It distinguishes itself from likely siblings like issue_invoice by explicitly stating that it mints no journal entry and stores no totals.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied through the tool's purpose: create an invoice from a sales order. The description does not explicitly state when to prefer this over issue_invoice or billing_generate_invoice, nor does it state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sales_order_listARead-only
Liste die Aufträge (P5): filter by status or contact. Each row carries the number, status, currency and order date. savedViewId support rides G00`s saved-view seam on the customization surface.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| contactId | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, covering the read-only behavior. The description adds the row fields (number, status, currency, order date), but the cryptic sentence about savedViewId ('rides G00`s saved-view seam') adds no clarity and may confuse. Minimal additional transparency beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, but the second sentence about savedViewId is confusing and unhelpful. The first sentence is clear, but the overall structure is not well-optimized for an agent; the cryptic phrase detracts from conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides the row fields, which is useful, but lacks details on pagination, sorting, or default behavior. It also omits explanation of workspaceId and does not clarify savedViewId. Adequate for a simple list tool, but incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains status and contactId via 'filter by status or contact', mentions savedViewId vaguely, but does not explain workspaceId (which is required) at all. Only half the parameters are meaningfully described.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it lists sales orders (Aufträge) and supports filtering by status or contact. The verb 'Liste' and the resource 'Aufträge' are explicit, and the tool is distinguished from siblings like sales_order_get (which retrieves a single order) and sales_order_create (which creates orders).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It specifies when to use the tool: to list orders with optional filters by status or contact. However, it does not explicitly contrast with alternatives like sales_order_get or mention when not to use it, so the guidance is clear but not fully exclusionary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_draftA
Create or update a draft journal entry (no money effect until posted).
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | ||
| date | Yes | ||
| lines | Yes | ||
| entryId | No | ||
| description | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that there is no money effect until posted, which is a key behavioral trait. However, it does not mention idempotency via idempotencyKey, the create-vs-update behavior based on entryId, or any side effects like validation failures.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that conveys the core purpose and the key distinction. It is front-loaded with the verb and avoids extraneous detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 7 parameters, a nested lines array, and no output schema, the description is insufficient. It omits parameter explanations, idempotency behavior, and return value expectations. It covers the core concept but leaves too much to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description gives no information about any of the 7 parameters. It does not explain what 'lines', 'idempotencyKey', or 'entryId' mean. The description provides no parameter semantics, leaving agents to guess.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb and resource: 'Create or update a draft journal entry'. The qualifier 'no money effect until posted' distinguishes it from posting tools like post_entry and reverse_entry, making its purpose unambiguous and differentiating it from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'no money effect until posted' implies this tool is for drafting and that posting tools are for finalizing. However, it does not explicitly name alternatives or state when to use create vs update. It provides contextual guidance but not explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_globalARead-only
Die globale Suche (global search): one query across every searchable record kind at once (Kontakte, Artikel, Belege, Projekte, Deals, Aufgaben, Lieferantenrechnungen, Aufträge, Bestellungen), including workspace-defined custom fields of type text/select/multiselect, findable the moment a value is set (no reindex exists). Results are grouped hits {entityKind, entityId, title, snippet?, matchedVia, route}, ranked exact > prefix > contains > custom-field-only (no relevance score is invented), deduped, paged by limit/cursor with hasMore/nextCursor. Every kind is fenced by that kind's own read capability: a kind the caller cannot read is silently absent, never counted, never hinted at, so an empty result set is indistinguishable from "no such record". q under 2 characters answers query_too_short; an entityKinds value outside the searchable roster answers unknown_entity_kind naming the roster; limit is capped at 50; savedViewId re-runs a saved global search (its stored q/entityKinds merge under any explicitly named field), so either q or savedViewId must be present. Reads only the local database and never local correspondence (mail, voice, drafts): those kinds are structurally absent from the roster. A pure read: writes nothing, needs no idempotencyKey.
| Name | Required | Description | Default |
|---|---|---|---|
| q | No | ||
| limit | No | ||
| cursor | No | ||
| entityKinds | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark readOnlyHint=true, and the description reinforces this with 'A pure read: writes nothing, needs no idempotencyKey.' It goes well beyond the annotation by disclosing result grouping, ranking semantics (exact > prefix > contains > custom-field-only), deduping, pagination, permission fencing (silently absent kinds), error conditions, and search freshness (no reindex). No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Though longer than typical descriptions, every clause delivers a necessary fact: result shape, ranking, dedup, pagination, permissions, errors, exclusions, and read-only guarantees. It is front-loaded with the core definition, and the density is justified for a cross-entity search tool with many behavioral edge cases.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is self-contained: it explains the return shape, paging contract, permission behavior, validation errors, parameter constraints, and scope limitations. Given there is no output schema and the parameter schema is minimal, this description thoroughly equips an agent to invoke the tool correctly and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden. It explains q (min length, query_too_short error), entityKinds (valid roster, unknown_entity_kind error), limit (capped at 50), savedViewId (re-runs saved search with merge behavior), and cursor (paging with hasMore/nextCursor). Every parameter gains meaning beyond its bare type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise definition: 'one query across every searchable record kind at once' and enumerates the entity kinds (Kontakte, Artikel, Belege, etc.). It clearly differentiates this global search from sibling kind-specific searches like serial_search, lot_search, files_search, and asset_search by emphasizing cross-kind coverage.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes clear context: this is the cross-entity search across all searchable records, with explicit exclusions (never mail, voice, drafts, which are 'structurally absent'). It does not name alternative tools for those exclusions, but the scope is stated so clearly that an agent can infer when to use it versus kind-specific searches.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_plugin_registryARead-only
Durchsuche das Erweiterungs-Verzeichnis (US-G02.5): query a configured registry through the OSS core`s registryClient seam and return a page of results (publisher, capability summary, requested permissions) that route into the same install-review flow as a local file. The OSS core wires NO client, so this answers needs_registry (never a hardcoded remote host); a configured registry that does not respond answers registry_unreachable (P9).
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | ||
| query | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true, the description goes well beyond the annotation by disclosing that no client is wired into the core, that no hardcoded remote host is ever used, that a non-responsive configured registry maps to registry_unreachable, and that results feed into the install-review flow. This is substantial behavioral context beyond what annotations already provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and mostly front-loaded, with purpose stated first and all sentences contributing meaningful behavior or error context. The mixed German/English phrasing and internal codes like US-G02.5 and P9 add slight noise but do not waste space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with no output schema, the description covers purpose, result content and error semantics, which is helpful. However, it lacks guidance on the required workspaceId, pagination behavior, and how the response is structured, leaving the agent with important gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It implies that 'query' is the search term and that results are paginated, but it never explains the required workspaceId, page format, or query semantics. An agent would have to infer too much from parameter names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: search the extension registry through the OSS core's registryClient seam and return a paged result set with publisher, capability summary, and requested permissions. It is clearly distinguishable from sibling tools like get_plugin_registry_entry by emphasizing directory search and paged results, though it does not explicitly name that sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete operational context: it is the tool to use when needs_registry applies because the OSS core wires no client, and it maps an unresponsive configured registry to registry_unreachable. It does not explicitly state when to prefer get_plugin_registry_entry or other plugin tools, but the intended flow is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_dunning_runA
Email an issued Mahnlauf's letters, one per debtor, through the configured transport. Settlement is re-checked against the OP-Liste (A16) at the moment of sending: a letter naming ANY invoice that CHANGED since issue (a payment partial OR full, a cancellation, or a credit note) is skipped whole, counted and named by its distinct cause (outcome paid | partially_paid | cancelled | credited, plus a changedItems list and the skippedSettled document ids on the result), because the frozen letter cannot shed an item and must never chase an invoice whose demand is no longer true (US-A15.4); the frozen run record stays untouched (D73). The same re-check rides get_dunning_run, so the Studio warns before a manual download too (the MIT core has no transport). Per-debtor outcomes are recorded (sent, no_email, send_failed, and the four change causes), a retry sends only what has not gone out, and the run turns 'sent' when every debtor's letter has. Degrades honestly: needs_email_config when no transport exists, needs_email_transport when one is configured but not wired, and a debtor without an email keeps a downloadable PDF. P8: outbound is draft-by-default, pass confirmed=true or enable the workspace dial. CONSEQUENCE: Hands the reminder letters to the outbound transport; sent reminders cannot be unsent.
| Name | Required | Description | Default |
|---|---|---|---|
| runId | Yes | ||
| confirmed | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without any annotations, this description carries the full burden of behavioral disclosure and exceeds expectations. It details the re-check against OP-Liste, the different change causes and their outcomes, skipped letter conditions, untouched run record, per-debtor outcome recording, retry semantics, degradation modes, draft-by-default requirement, and irreversibility. This is a model of transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place. It is front-loaded with the core action, then systematically covers re-check logic, outcomes, degradation, confirmation, and consequences. No filler or redundancy. The length is appropriate for the complexity of the operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no annotations and no output schema, the description provides nearly everything needed to call the tool correctly: preconditions, behavior, outcomes, edge cases, required confirmation, and non-reversibility. It even documents the partial result structure (cause, changedItems, skippedSettled). Only a formal output schema is missing, but that is beyond the description's responsibility.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains the 'confirmed' parameter ('pass confirmed=true or enable the workspace dial') and implies the roles of runId/workspaceId/idempotencyKey through context. IdempotencyKey is standard and not elaborated, but the named parameters are clear enough. A 4 is justified because the description adds meaningful parameter context beyond parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Email an issued Mahnlauf's letters, one per debtor, through the configured transport.' It clearly distinguishes this from sibling tools like propose_dunning_run, issue_dunning_run, and get_dunning_run by focusing on the sending action and the settlement re-check. No ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use the tool (after a run is issued, to send letters) and mentions retry behavior ('a retry sends only what has not gone out'). It also notes that the same re-check applies to get_dunning_run, which helps an agent understand when the send path vs. preview path is warranted. However, it does not explicitly name sibling alternatives or state exclusions, so the guidance is slightly below par.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_invoiceB
Send an issued invoice by email (renders + attaches the PDF, transitions to sent). Outbound step is P8-gated: pass confirmed=true or enable the workspace dial. CONSEQUENCE: Hands the invoice to the outbound transport; a sent invoice cannot be unsent.
| Name | Required | Description | Default |
|---|---|---|---|
| No | |||
| confirmed | No | ||
| invoiceId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden for behavioral disclosure. It explicitly warns of an irreversible side effect ('a sent invoice cannot be unsent'), explains the outbound handoff, and describes the rendering/attachment behavior. The unexplained 'P8-gated' jargon and workspace dial prevent a perfect score.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, with each sentence adding meaningful content: action, gate, and irreversible consequence. No fluff is present, though the cryptic 'P8-gated' and 'workspace dial' phrasing could be clearer without much additional length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there are no annotations, no output schema, and five parameters with zero schema descriptions, the description is not complete enough. It omits semantics for the optional `email` and `idempotencyKey`, and does not describe what the caller should expect in return. The core action and risk warning are present, but important invocation details are missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only explains the `confirmed` parameter. The meaning and purpose of `email`, `idempotencyKey`, `workspaceId`, and `invoiceId` are not described beyond the generic action of sending an invoice. This leaves most parameters under-specified for an agent invoking the tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action: 'Send an issued invoice by email' with explicit outcomes ('renders + attaches the PDF, transitions to sent'). It implies the tool operates on already-issued invoices, which loosely distinguishes it from sibling issue_invoice, but it does not name any alternative. This is clear but not fully differentiated from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives practical gate guidance ('pass confirmed=true or enable the workspace dial') and implies the invoice must already be issued. However, it does not explicitly state when to choose this tool over alternatives like issue_invoice or billing_generate_invoice, nor does it state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
serial_archiveA
Soft-archive a serial (status archived). Refused with serial_has_balance while the unit is still available or reserved (issue, return or scrap it first). Deletion is never offered.
| Name | Required | Description | Default |
|---|---|---|---|
| serialId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses that this is a soft archive (status change) rather than a delete, and it names the specific refusal error serial_has_balance with the condition that triggers it. It does not mention reversibility or response behavior, but the key mutation semantics are clearly stated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The entire description is two sentences and every sentence earns its place: the first states the core action, the second gives the key refusal condition and the 'deletion is never offered' clarification. The message is front-loaded and compact with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers the primary failure mode and the soft-archive nature, but it omits several important operational details: what the idempotencyKey is for, what a successful call returns, and how statuses like 'available' or 'reserved' are determined. The description is solid but not fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description adds almost no parameter-level information. The text mentions 'serial' only through the tool name, and there is no explanation of serialId, workspaceId, or idempotencyKey semantics. Since the schema does not compensate and the description does not fill the gap, parameter meaning is left largely to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('soft-archive') and a specific resource ('serial'), and clearly distinguishes the behavior from deletion with 'Deletion is never offered.' It also clarifies the resulting status ('status archived'), making the tool's purpose unambiguous even among many sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: it explains that the operation will be refused with serial_has_balance while the unit is available or reserved, and advises issue, return, or scrap first. It does not explicitly name alternative sibling tools (e.g., serial_set_status), which would have made the when-to-use guidance fully explicit, but the exclusion condition effectively tells the agent when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
serial_createA
Create one serial (an individually identified unit) against a serial-tracked item. number is 1-60 chars, unique per item case-insensitively (serial_number_taken); the item must be serial-tracked (else tracking_not_applicable). When the item is lot_and_serial a lotId is required (lot_reference_required) and must belong to the same item (lot_item_mismatch). Starts status available.
| Name | Required | Description | Default |
|---|---|---|---|
| lotId | No | ||
| notes | No | ||
| itemId | Yes | ||
| number | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It transparently describes the initial status ('available'), the uniqueness requirement (case-insensitive), and specific error conditions (serial_number_taken, tracking_not_applicable, lot_reference_required, lot_item_mismatch). The description does not disclose side effects beyond creation, but it provides essential behavioral context for an agent to anticipate failures and outcomes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, with three dense sentences that front-load the core purpose and then provide necessary constraints. Every sentence adds valuable information without redundancy. The structure is organized logically from definition to specific conditions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (6 parameters, no output schema), the description covers the most critical preconditions and error cases, which is sufficient for an agent to decide when to call it. It does not explain the return value or any additional side effects beyond status, but that may be acceptable given the lack of output schema and the focus on creation. The description compensates for the 0% schema coverage on parameters, but leaves some parameters (idempotencyKey, workspaceId) undocumented, which is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
While schema description coverage is 0% and the schema provides no descriptions for parameters, the description adds significant meaning to the 'lotId' parameter (required for lot_and_serial items and must match the item), the 'number' parameter (constraints on length and uniqueness), and the 'itemId' parameter (must be serial-tracked). It does not explicitly explain 'workspaceId', 'idempotencyKey', or 'notes', but the essential parameters are clarified. The addition of error conditions helps an agent understand parameter relationships.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to create a single serial against a serial-tracked item. It specifies the exact resource (serial), the action (create), and key constraints (item must be serial-tracked). It also distinguishes itself from sibling tools like serial_create_bulk by focusing on single serial creation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context on when the tool is applicable: for serial-tracked items, and when a lot is required (for lot_and_serial items). However, it does not explicitly state when to use serial_create_bulk instead, though the name implies bulk creation. It gives no exclusions beyond the mentioned error conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
serial_create_bulkA
Create many serials at once for one item from a list of numbers (receiving a carton). ALL-OR-NOTHING: any duplicate number (against the store or within the batch) aborts the whole call with serial_number_taken and writes nothing. Same lotId rules as serial_create.
| Name | Required | Description | Default |
|---|---|---|---|
| lotId | No | ||
| itemId | Yes | ||
| numbers | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it discloses critical behavior: all-or-nothing atomicity, duplicate detection against the store and within the batch, the resulting error code serial_number_taken, and that nothing is written on failure. It leaves idempotency semantics and success-response shape unspecified, but the essential safety/critical failure behavior is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: it states the action, the resource, the batch context, and the critical atomic/duplicate behavior in two clear sentences. Every sentence adds value and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 5 parameters, no output schema, and no annotations, the description does more than the minimum by explaining the batch context and atomic failure semantics. It is not fully complete because it does not define success behavior, idempotencyKey usage, or the referenced lotId rules, but it is sufficient for an agent to understand the core call behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it partially does by explaining that numbers is the list being created and that lotId follows serial_create's rules. However, workspaceId, itemId, and idempotencyKey remain semantically unexplained, and the lotId reference depends on knowledge of serial_create.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('create many serials at once for one item from a list of numbers') with a clear scenario ('receiving a carton'), and the bulk nature distinguishes it from the sibling serial_create. This is a precise verb-resource-scope definition, not a tautology.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly indicates this is for batch-creating serials from a list of numbers, gives a concrete receiving-carton context, and cross-references serial_create for lotId rules. It does not explicitly state when-not-to-use or name alternatives, but the bulk vs. single distinction is strongly implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
serial_getARead-only
Read one serial by id, including its lot link, status and current-location projection.
| Name | Required | Description | Default |
|---|---|---|---|
| serialId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already exists, and the description reinforces it with 'Read' while adding useful output context (lot link, status, current-location projection). It does not cover error or not-found behavior, but for a read-only getter this is a reasonable level of transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence states the action, target, and key returned fields with no filler. Every part of the description contributes meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-string getter with no output schema, the description gives sufficient context: what it reads, how it is keyed, and what key data it returns. Minor omissions like not-found or error behavior are acceptable given the tool's simplicity and read-only annotation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies that serialId is the lookup key ('by id'), but it does not explain workspaceId's role or expected value format; the parameter names are self-explanatory enough to avoid major ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific operation, 'Read one serial by id', and names specific returned aspects (lot link, status, current-location projection). This clearly distinguishes it from serial_list and serial_search, which are sibling tools for finding or listing serials.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'Read one serial by id' makes it clear this is for single-record retrieval by identifier, implying it should be used over list/search alternatives. It does not explicitly name alternatives or state when not to use it, but the use case is clear enough for a simple getter.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
serial_listARead-only
List serials, filterable by itemId, lotId, current locationId, status, and a case-insensitive search over number. Archived serials are hidden unless includeArchived or an explicit status is given. Accepts a savedViewId (G00 saved-view seam).
| Name | Required | Description | Default |
|---|---|---|---|
| lotId | No | ||
| itemId | No | ||
| search | No | ||
| status | No | ||
| locationId | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| includeArchived | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already supply readOnlyHint=true, and the description adds behavioral detail beyond that: archived serials are hidden unless includeArchived or an explicit status is given, search is case-insensitive over number, and savedViewId follows the G00 seam. This helps an agent predict results without contradicting the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three compact sentences with no filler. It front-loads the core list/filter operation, then adds archiving behavior and the saved-view hook, all in under 30 words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool, the description covers the operation, filters, archiving semantics, and the savedViewId seam, which is enough for an agent to invoke it correctly. Pagination and return-shape details are absent, but this is a modest gap given the simple list purpose and readOnly annotation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the bare schema. It meaningfully explains most parameters: itemId, lotId, current locationId, status, case-insensitive search over number, includeArchived, and savedViewId. Only workspaceId, the required parameter, is not explicitly described, but its role is reasonably inferable from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (List) and resource (serials), and enumerates the main filter dimensions. It is clearly a list/query operation and easy to tell apart from single-record siblings like serial_get, though it does not explicitly contrast with serial_search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use this tool: any listing of serials filtered by item, lot, location, status, or search term. It documents the archived-serial behavior and savedViewId seam, which informs invocation. It does not state exclusions or name alternatives, but the intended use is obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
serial_searchBRead-only
Full-text serial search over number (case-insensitive), capped at 100 rows. Empty query returns an empty list.
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true, so the description adds valuable behavior details: case-insensitivity, a 100-row cap, and that an empty query returns an empty list. These go beyond what annotations provide, giving the agent expectations about result limits and empty-input handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence that front-loads the purpose and key behavioral constraints. It contains no filler and every clause adds meaningful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and only a readOnly annotation, the description should cover purpose, parameters, and behavior. It covers purpose and some behavior but omits parameter semantics, especially workspaceId, and gives no hint of the return format or pagination. An agent would not know how to construct a valid call beyond guessing the query parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for both parameters. It implies 'query' is the search term via 'over number', but it does not explain 'workspaceId' at all, nor does it clarify whether query is optional (schema shows it is not required). This leaves a required parameter undocumented and the query format ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool performs a full-text search over serial numbers with a case-insensitive match. It identifies the resource (serial numbers) and the action (search), distinguishing it from serial_get (exact lookup) and serial_list (listing). However, it does not explicitly contrast with sibling search tools like lot_search, so it is not perfectly differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention serial_list or serial_get, nor any conditions that would select this tool over them. An agent must infer from the name and description that it is for searching by number.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
serial_set_statusA
Change a serial status: available | reserved | issued | returned | scrapped | archived, with an optional reason. Scrapping an available unit is allowed and removes it from availability. Never mints a stock movement.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| status | Yes | ||
| serialId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that scrapping an available unit is allowed and removes it from availability, and explicitly states that no stock movement is created. This is important behavioral context that helps an agent understand side effects and limitations, though it could go further by noting whether status changes are reversible or what happens to reserved/issued items.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise (one sentence), front-loaded with the core action and allowed values, and every clause adds value. It includes the important edge case of scrapping and the critical exclusion about stock movements. There is no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (5 params, 4 required) and no output schema, the description covers the essential behavioral aspects: the allowed statuses, the scrapping rule, and the stock movement exclusion. It could be more complete by specifying how the status affects related entities (e.g., reservations) or whether the tool validates transitions, but the current information is sufficient for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions the 'reason' parameter as 'optional reason' and lists the allowed values for 'status', but does not describe the semantics of 'workspaceId', 'serialId', or 'idempotencyKey'. Those are likely self-explanatory from their names, but the description could add more detail on how these parameters interact, e.g., whether idempotencyKey prevents duplicate status changes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Change') and resource ('serial status'), lists all allowed status values, and mentions an optional reason parameter. It also adds important nuance by clarifying that scrapping an available unit is allowed and removes it from availability, which distinguishes it from generic status updates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when changing a serial's status. It doesn't explicitly describe when not to use it, but the context of sibling tools like serial_update and serial_archive provides implicit contrast. The phrase 'Never mints a stock movement' clarifies that it should not be used for inventory movements that affect stock, which is useful guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
serial_updateA
Edit a serial through a descriptive patch (number if still unique, notes, lotId which must belong to the same item). The status is changed via serial_set_status, not here. Only the fields present in patch change.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | No | ||
| serialId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It covers meaningful constraints: partial-update behavior, uniqueness of the number, and the cross-entity invariant that lotId must belong to the same item. It does not mention idempotencyKey semantics, possible failure cases (e.g., duplicate number), or return value, but it still provides essential behavioral context beyond the bare schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three crisp sentences, front-loaded with the core purpose and field list. Every clause contributes: patch fields, the status boundary, the lotId rule, and the partial-update guarantee. No filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mid-complexity mutation tool with no output schema and no annotations, the description covers the key facts an agent needs: what can be edited, the uniqueness constraint, the lotId relationship, and the explicit routing of status changes elsewhere. It leaves out idempotencyKey/workspaceId semantics, but those are conventional and the schema names make them self-obvious. Slightly more detail about expected errors would push it to a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds real semantic value to the patch fields: number 'if still unique', lotId 'must belong to the same item', and the note that only present fields change. The top-level parameters, serialId/workspaceId/idempotencyKey are not described, but they are not particularly ambiguous given their names; still, the description focuses on patch semantics and leaves the rest undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Edit a serial') and enumerates the editable fields (number, notes, lotId). It explicitly contrasts with serial_set_status, making the tool's scope unambiguous and distinct from the sibling that handles status changes. This is a clear, differentiated purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says status changes belong in serial_set_status and not here, giving a clear 'when not to use this tool' signal. It also clarifies the partial-update semantics ('Only the fields present in patch change'). However, it does not mention other sibling alternatives like serial_create or serial_archive, though those are less related to editing an existing serial.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_agent_dialA
Set one approval-dial level to ask or auto for one governed capability (post/issue/send/dun/pay/vat-file/customize/plugin-install/close-period/go-live). ask drafts the agent write for human approval; auto executes it when RBAC also permits. An unknown capability or level is rejected. Owner-only (manage_agent_dial).
| Name | Required | Description | Default |
|---|---|---|---|
| level | Yes | ||
| capability | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It discloses side effects ('ask drafts... auto executes'), the RBAC condition, rejection of unknown inputs, and the owner-only permission. This is strong coverage, though it does not explain what happens with repeated calls despite the idempotencyKey parameter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no filler. The main action and key scope come first, followed by the ask/auto semantics, validation behavior, and permission requirement. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation with four required parameters, no annotations, and no output schema, the description covers the essential operational context: allowed values, behavior, validation, and authorization. It is slightly incomplete on the role of idempotencyKey and does not mention the sibling get_agent_dial, but overall an agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% parameter coverage and no enum constraints, so the description must add meaning. It does clarify meaningful values for capability and level, but leaves workspaceId and idempotencyKey to be inferred from their names and from the general context, which is partially compensating but incomplete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Set one approval-dial level') and clearly names the resource ('to ask or auto for one governed capability') and enumerates the exact capability values. This makes it easy to distinguish from sibling tools like get_agent_dial and approve_drafted_action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for what the tool does and even explains the behavioral difference between ask and auto, plus the owner-only permission. However, it never explicitly tells the agent when to choose this tool instead of siblings like get_agent_dial for reading the dial or approve_drafted_action for handling a drafted action.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_aging_bucket_configA
Redefine where the aging buckets cut, as a strictly increasing list of positive day counts (the default [30, 60, 90] gives 0-30, 31-60, 61-90 and 90+). This changes the VIEW only: the buckets re-partition the same receivables total and can never change it, and the reconciliation to 1100 Debitoren holds for any boundary set. It is a DACH reporting convention, not a statutory figure, and it is not a dunning deadline (A15 owns those). A repeat call under the same idempotencyKey writes nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| boundariesDays | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description discloses non-destructive behavior ('changes the VIEW only', never changes totals), confirms reconciliation integrity, and specifies idempotency (repeat calls write nothing). This fully covers the safety profile and side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each adding value: purpose+example, scope clarification, and idempotency note. Front-loaded with the core action, no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a config setter with three parameters and no output schema, the description covers the purpose, constraints, exclusions, and behavioral guarantees. Nothing essential is missing for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description compensates: it explains boundariesDays format (strictly increasing positive day counts) with an example, and explains idempotencyKey semantics. workspaceId is left to convention, but that is standard and acceptable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('redefine') and resource ('aging bucket config'), and clarifies the effect with an example. It also distinguishes from dunning deadlines and statutory figures, so an agent can identify its scope without confusion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly states it is a view-only setting for DACH reporting and explicitly excludes dunning (A15 owns those). It does not name sibling tools like get_aging_bucket_config or aging_report, but the scope is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_bank_opening_balanceA
Post a Bankkonto's opening balance as a real balanced journal entry (debit the bank account, credit 9100 Eröffnungsbilanz), in integer Rappen. Explicitly confirmed and idempotent: re-sending the same idempotencyKey returns the original entry and never posts twice. A correction to a posted opening balance is a reversing entry, never a second opening.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | ||
| fxRate | No | ||
| currency | No | ||
| amountMinor | Yes | ||
| workspaceId | Yes | ||
| bankAccountId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It covers the accounting effect (balanced journal entry, debit/credit accounts), the integer Rappen constraint, idempotency behavior, and correction semantics. It does not mention prerequisites, output details, or broader side effects, but the key behavioral traits are transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, well-structured, and front-loaded. Each sentence adds essential information: what is posted, the accounting entry, the integer Rappen constraint, idempotency, and how corrections should be handled. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no output schema, the description covers the most important behavioral context and idempotency semantics well. It is incomplete regarding the meaning of fxRate and currency, prerequisites for the bank account, and what the response contains, leaving notable gaps for an agent to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds important meaning for amountMinor (integer Rappen) and idempotencyKey (safe retry semantics), and implies bankAccountId via the posting context. However, workspaceId, date, currency, and fxRate are not explained, and schema description coverage is 0%, so the description only partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action: posting a Bankkonto's opening balance as a real balanced journal entry with the exact debit and credit accounts. It also distinguishes itself from related tools like preview_bank_opening_balance by emphasizing that this tool actually posts and is idempotent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear operational guidance on idempotency: re-sending the same idempotencyKey returns the original entry and never posts twice. It also tells users that corrections must be reversing entries, not a second opening. It does not explicitly name alternatives or exclusion conditions, but the guidance is practical and clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_bank_sync_scheduleA
Link or unlink the G01 automation rule that drives this EBICS connection's scheduled sync (A36 US-A36.3). The channel panel's Automatischer Abruf toggle creates a schedule rule whose action is bank_sync for this connectionId through G01's rule verbs, then calls this to store the linkage on the connection so the status card shows the cadence and a retire clears it. Pass ruleId to link (the rule must have action bank_sync and not be archived), or omit/null to unlink. Creates, enables or fires NO rule (that is G01's surface); default is OFF. Idempotent; banking-write. Posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| ruleId | No | ||
| workspaceId | Yes | ||
| connectionId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so richly: it discloses idempotency, banking-write semantics, default OFF, that it creates/enables/fires no rule, that it stores linkage on the connection, and that it posts nothing. This goes well beyond a typical definition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The definition is densely informative but front-loaded: the core purpose comes first, followed by the mechanism, invocation rules, exclusions, and side-effect profile. Every sentence earns its place and there is little redundant prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is unusually complete given no annotations and no output schema: it covers purpose, preconditions, defaults, side effects, and the boundary with G01. The main gap is that the semantics of the required workspaceId and idempotencyKey parameters are left to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaningful semantics for ruleId (link with preconditions, omit/null to unlink) and references connectionId as the linkage target. However, schema coverage is 0% and required params workspaceId and idempotencyKey are not explained, so the description only partially compensates for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence names the exact verb pair (link or unlink), the resource (G01 automation rule driving an EBICS scheduled sync), and the scope. It also explicitly scopes out rule creation, enabling, or firing, which distinguishes it from bank_sync and rule-management siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives direct invocation guidance: pass ruleId to link, omit/null to unlink, and requires the rule to have action bank_sync and not be archived. It also explains that creating/enabling/firing rules belongs to G01's surface, preventing misuse of this tool for those operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_camt_matchingA
Tune the workspace debit-matching (A36 §6b): valueDateWindowDays (0-60, default 5) widens or narrows the value-date proximity signal, and reviewThreshold (any|medium|high, default any) sets the confidence floor below which a booked debit reports needsReview. These tune ranking and the review event only; they never touch the mandatory exact amount + currency gate. Idempotent; one row per workspace.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| reviewThreshold | No | ||
| valueDateWindowDays | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden. It discloses idempotency ('Idempotent; one row per workspace'), parameter constraints (ranges and defaults), and explicitly scopes side effects to ranking and review events only. This gives an agent a clear model of the tool's behavior beyond what a bare schema would provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the purpose, and each clause earns its place. It packs parameter semantics, scope limitations, and idempotency without redundancy. This is a model of concise, high-density documentation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple configuration tool with no output schema operative, the description covers all parameters and key behavioral aspects. It does not mention prerequisites, permissions, or return values, but given the limited scope and idempotency, an agent can safely invoke it with the provided information. The one-row-per-workspace constraint is a useful caveat.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains both optional parameters: valueDateWindowDays (0-60, default 5) and reviewThreshold (any|medium|high, default any). The required parameters workspaceId and idempotencyKey are not described, but they are conventional and self-explanatory. The description adds meaningful constraints and defaults that the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Tune the workspace debit-matching (A36 §6b)'. It clearly distinguishes this from operational matching tools like suggest_matches or confirm_match by focusing on configuration parameters. The phrase 'tune ranking and the review event only' further clarifies the narrow scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for adjusting matching thresholds but does not explicitly state when to choose this over alternatives. It mentions what it does not affect (the exact amount + currency gate), which provides some context, but no sibling tools are named and no 'when-to-use' vs 'when-not-to-use' guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_creditor_bank_profileA
Set (or correct) the IBAN a vendor is paid at. A17's vendor bill carries no IBAN of its own; this is where a creditor's bank details live so create_payment_batch can build a pain.001 instruction. Validates the IBAN (A19's validateIban) and derives whether it is a QR-IBAN, which decides whether the bill's own vendorReference must be a QRR (a QR-IBAN never accepts free-text remittance). This alters who a future payment reaches: it is on the automation denylist (D65 leg f), mirroring update_bank_account and set_creditor_profile.
| Name | Required | Description | Default |
|---|---|---|---|
| iban | Yes | ||
| vendorId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of explaining side effects. It clearly states that the tool alters who a future payment reaches, validates the IBAN via a known validator, derives QR-IBAN status, and imposes constraints on vendorReference. This is strong disclosure of behavioral consequences beyond a simple setter. It could mention failure behavior or idempotency, but the provided detail is well above the minimum.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening sentence is a strong, front-loaded summary. However, the rest is dense and references internal identifiers (A17, A19, D65) that may be opaque to a caller. Every sentence adds useful integration context, but the overall structure is not as tight as it could be; the internal references could have been paraphrased for clearer standalone value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a mutation tool with no output schema and no annotations, yet the description explains its role in the payment-batch pipeline, validation behavior, QR-IBAN consequences, and automation denylist status. It leaves some gaps, such as return value, idempotencyKey behavior, and parameter-level details, but given the richness of the rest of the context, an agent has enough to call it correctly in most situations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not systematically explain the four parameters. It adds meaning for 'iban' (validated, QR-IBAN derivation, payment routing implications) and 'vendor' indirectly, but does not explain workspaceId, vendorId semantics, or idempotencyKey. Since the schema provides no descriptions, the description was expected to compensate but only partially covers iban.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and object: 'Set (or correct) the IBAN a vendor is paid at.' It explains the role of this tool in relation to vendor bills, create_payment_batch, and pain.001, making the tool's purpose unambiguous. It also subtly distinguishes itself from related sibling tools like update_bank_account and set_creditor_profile by naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use this tool: when a vendor's IBAN must be set or corrected, noting that vendor bills carry no IBAN and that this is the place where creditor bank details live for payment creation. It does not explicitly list when to use an alternative instead, but it does reference related tools (update_bank_account, set_creditor_profile) as mirrors, which gives some differentiation. Missing explicit exclusion criteria, but enough guidance for an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_creditor_profileA
Set the QR-bill creditor: IBAN (a QR-IBAN gives a QRR reference, a plain IBAN gives SCOR), optional name (defaults to the company name) and optional structured address. The IBAN saves without the address; the address is needed before the first QR-bill renders (buildQrBill answers needs_creditor_address until then). A partial address is refused (needs_structured_address).
| Name | Required | Description | Default |
|---|---|---|---|
| iban | No | ||
| qrIban | No | ||
| address | No | ||
| workspaceId | Yes | ||
| creditorName | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and discloses key behaviors: the IBAN saves without the address, the address is required before rendering, and partial addresses are refused. It also explains the QRR/SCOR mapping. This is strong behavioral transparency, though it omits aspects like idempotency or return values.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three compact sentences convey the main action, parameter nuances, and validation behavior without fluff. The most important information is front-loaded, and each sentence adds meaningful content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters, a nested object, no output schema, and no annotations, the description covers the essential workflow (address requirement, buildQrBill integration) but leaves gaps: 'qrIban' semantics, the role of 'workspaceId', and what the tool returns. It is useful but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies the 'iban' semantics (QR-IBAN vs plain), the 'creditorName' default, and that 'address' must be complete. However, it does not explain the 'qrIban' parameter at all, and 'workspaceId' is only known from the schema. This is partial compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the action ('Set the QR-bill creditor') and the resource (creditor profile), and adds differentiating detail about IBAN types (QR-IBAN vs plain) and reference generation (QRR vs SCOR). This distinguishes it from related tools like set_creditor_bank_profile.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage context: the address is needed before the first QR-bill renders, referencing buildQrBill. However, it does not explicitly state when to use this tool versus alternatives (e.g., set_creditor_bank_profile), nor provide exclusions or preferred conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_default_document_templateA
Make a template THE default for its document kind: the prior default clears in the same atomic write (at most one default per kind), and every document of that kind issued afterwards freezes to it. Wide-effect and reviewable: this changes the workspace's outward face.
| Name | Required | Description | Default |
|---|---|---|---|
| templateId | Yes | ||
| workspaceId | Yes | ||
| documentKind | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses important behavioral traits: the prior default clears in the same atomic write, at most one default per kind, and every document issued afterwards freezes to the new default. It also notes the wide-effect and reviewable nature of the change. This goes beyond the annotations (which are absent) and provides critical context for an agent deciding whether to invoke this tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded. The core action is stated first, followed by the key behavioral implications. Every sentence adds value, and the description is appropriately sized for the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential behavioral context: the atomic clearing of the prior default, the one-default-per-kind constraint, and the wide-effect nature. It doesn't mention return values, but there is no output schema, so the description could be expected to explain what the agent should expect. However, the behavioral context is strong enough that an agent can likely infer the outcome. A small gap remains regarding the response format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. It mentions 'document kind' and 'template' but doesn't explicitly map to the parameters (workspaceId, documentKind, templateId, idempotencyKey). The description adds some meaning by explaining the relationship between documentKind and templateId, but it doesn't clarify the role of workspaceId or idempotencyKey. Baseline 3 is appropriate given the partial compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: making a template the default for its document kind. It specifies the resource (document template), the action (set as default), and the scope (per document kind). It also distinguishes itself from sibling tools like create_document_template, update_document_template, and archive_document_template by focusing on the default-setting behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when a template should become the default for its document kind. It doesn't explicitly name alternatives or exclusions, but the context of 'default' and 'document kind' makes the use case clear. It could be improved by explicitly stating when not to use it (e.g., for creating or updating templates).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_diagnosticsA
Turn error recording on or off for this computer. Off is the default, and turning it off also deletes what was recorded. Set this only when the person asks for it.
| Name | Required | Description | Default |
|---|---|---|---|
| capture | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose behavioral traits. It does so by noting that off is the default, that turning it off deletes recorded data (a destructive side effect), and that it should only be set on explicit request. This is valuable transparency beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with zero filler. It front-loads the core action and includes the critical side effect and usage condition, making it highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is a simple toggle, and the description covers the action, side effect, and when to use it. However, the workspaceId parameter remains unexplained, which is a notable gap for a required parameter. Without an output schema, the description could also mention expected return, but for a setter this is often not necessary. Overall, adequate but with a clear omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no descriptions for the parameters (0% coverage). The description implicitly explains 'capture' as the on/off toggle but says nothing about 'workspaceId'. Since workspaceId is required and undocumented in both schema and description, the description fails to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Turn error recording on or off') and the resource ('this computer'), which distinguishes it from sibling tools like get_diagnostics (retrieve) and clear_diagnostics (clear). It also includes the default state and the deletion side effect, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear usage condition: 'Set this only when the person asks for it.' This tells the agent not to invoke it proactively. However, it does not explicitly mention alternatives like get_diagnostics or clear_diagnostics, so the guidance is not fully exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_dunning_configA
Set the Mahnwesen policy: all three levels in one write, thresholds strictly increasing. A Mahngebühr has NO statutory basis and is chargeable only when contractually agreed, so it defaults to zero; a positive fee always BOOKS (the letter demands exactly what books) and needs an income account. There is no tax code to configure: the fee's VAT splits pro rata across the chased invoice's own rate bases at issue (D69, ESTV practice: the fee is part of the underlying supply's Entgelt). minIntervalDays per level is the minimum days that must pass after the previous level's letter issued before this level is reached (default 10), so an invoice already past every threshold escalates one letter per interval rather than 1./2./3. in three days; 0 disables the spacing for that level. The Verzugszins rate floor is 500 bp (Art. 104 OR) and the note is computed on the invoice principal, actual days over 365. A repeat call under the same idempotencyKey writes nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| levels | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so exceptionally well. It discloses major behavioral traits: fee defaults to zero, positive fees book immediately and require an income account, VAT handling, minIntervalDays escalation semantics, the interest-rate floor, and idempotency write-nothing behavior. This exceeds what a generic 'set config' description would provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence front-loads the tool's purpose and core constraint. The remaining sentences are dense but each adds a distinct behavioral, legal, or accounting rule. There is no filler or repetition of schema-visible names.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a complex configuration tool with no annotations and no output schema, and the description covers nearly all important caveats: bookkeeping, VAT, intervals, interest, and idempotency. It stops just short of complete by not explaining a few nested fields such as templateKey and showInterest, or the success/result shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It adds deep meaning for levels, minIntervalDays, fee fields, income account requirements, interest basis, and idempotencyKey. It does not explicitly map a few nested fields like templateKey or showInterest, so it is not a perfect 5, but it adds substantial value beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Set the Mahnwesen policy'), the scope ('all three levels in one write'), and a key invariant ('thresholds strictly increasing'). It differentiates this from the dunning sibling tools by emphasizing this is the policy setter, not a run/propose operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for what this tool is used for and its atomic, all-levels behavior. However, it never explicitly contrasts it with alternatives such as get_dunning_config, propose_dunning_run, or issue_dunning_run, leaving the when-not-to-use guidance implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_ebill_configA
Record the eBill-Biller-ID after enrolling with a certified network partner (a commercial step outside the product). Validates the SWP billerPid shape (41 followed by 15 digits) and upserts the one config row per workspace; a malformed id returns invalid_biller_pid and persists nothing. Naturally idempotent (asserts an absolute state), so it takes no idempotency key. Posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| billerPid | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the validation behavior, the upsert semantics, the idempotency rationale, and the fact that it posts nothing. This is unusually transparent for a config setter.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each adding distinct value: the prerequisite, the validation, the persistence semantics, and the idempotency note. No filler, no repetition of schema field names beyond what's needed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter config setter with no output schema, the description covers the prerequisite, the validation rule, the persistence behavior, and the idempotency design. An agent has everything needed to decide when to call it and what to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the billerPid shape ('41 followed by 15 digits') and the workspace scoping ('one config row per workspace'). It doesn't explicitly define workspaceId, but the context makes it clear enough.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Record'), a precise resource ('eBill-Biller-ID'), and the exact context (after enrolling with a certified network partner). It also names the sibling get_ebill_config implicitly by describing the config row per workspace, and the validation rule distinguishes it from generic config setters.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use it ('after enrolling with a certified network partner'), what it does not do ('Posts nothing'), and what happens on invalid input ('returns invalid_biller_pid and persists nothing'). It also explains why no idempotency key is needed, which is a clear usage guideline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_field_valueA
Set one custom field value on one record, validated against the field type (a money value is an integer Rappen count, a date is ISO-8601, a select value is one of its options). Passing null clears the value. Requires whatever capability editing that entity itself requires, never a G00 capability: a custom field is not a side door.
| Name | Required | Description | Default |
|---|---|---|---|
| value | No | ||
| entityId | Yes | ||
| fieldKey | Yes | ||
| entityKind | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does well by disclosing type-validation behavior (with concrete examples), the null-clearing side effect, and the capability requirement. This goes well beyond a generic 'sets a field' statement and warns against privileged misuse. It doesn't mention idempotency or error behavior, but the core hidden behaviors are surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two packed sentences with no filler. The primary action is stated first, followed by critical validation details, then the null behavior, then the security caveat. Every clause earns its place, and the structure guides the reader from what to how to when-not-to.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (multiple value types, null clearing, capability policy) and the absence of annotations/output schema, this description covers the most decision-critical aspects. It lacks explicit notes about idempotency and response shape, but those are less likely to cause misuse. For a mutation tool, it is notably more complete than typical.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It meaningfully explains the 'value' parameter (formats for money, date, select, and null semantics), which is the most complex parameter. However, it does not clarify the roles of entityKind, fieldKey, idempotencyKey, workspaceId, or entityId beyond what their names imply. The added value is partial but relevant.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Set'), a precise resource ('one custom field value on one record'), and adds essential constraints (validated by type, null clears). This clearly differentiates it from unrelated sibling operations and would let an agent understand exactly what the tool does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear context for use (setting a custom field with type validation and clearing via null), and includes an important caveat about required capabilities ('requires whatever capability editing that entity itself requires, never a G00 capability'). It does not explicitly name alternatives or say when not to use this tool, but the context is strong enough that an agent would not confuse it with list/define operations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_fiscal_configA
Set legal form, base currency, or fiscal year start (currency/year lock once the ledger is non-empty).
| Name | Required | Description | Default |
|---|---|---|---|
| legalForm | No | ||
| workspaceId | Yes | ||
| baseCurrency | No | ||
| fiscalYearStart | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the lock behavior: 'currency/year lock once the ledger is non-empty', which is a critical behavioral trait. It also implies that setting these fields may be restricted by ledger state, which is valuable. However, it doesn't mention whether legal form has similar restrictions or reversal possibilities, but the explicit lock mention is a strong plus.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that packs the purpose and the key constraint (lock) effectively. It is concise, with no fluff, and the lock behavior is highlighted, making it front-loaded. Every part adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 4 parameters (with 1 required) but no output schema, the description is moderately complete. It covers the key settings and the lock caveat, but lacks details on parameter formats, default behaviors, and possibly side effects. For an agent to call it correctly, it would need to infer date/currency formats. The lock behavior is well disclosed, but the complexity of fiscal configuration suggests more could be said.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. The description lists the three settings (legalForm, baseCurrency, fiscalYearStart) but does not provide syntax or format details (e.g., date format for fiscalYearStart, currency codes). It omits workspaceId, which is required. Still, it adds meaning by linking parameters to the operation, but not fully detailed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states that the tool sets legal form, base currency, or fiscal year start, which are concrete resources/attributes. It distinguishes itself from sibling tools like set_vat_method or set_fx_method by specifying these particular fields, though it doesn't explicitly name alternatives. It's not a tautology and clearly indicates a set operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool (when configuring fiscal settings) but does not provide explicit guidance on when not to use it or mention alternative tools. For example, it doesn't distinguish from set_vat_method or set_fx_method, which could be relevant. The context is clear but lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_fx_methodA
Elect the MWSTV Art. 45 conversion basis for a Steuerperiode: daily (Tageskurs, Devisenkurs Verkauf), monthly_avg (Monatsmittelkurs) or group (Konzernumrechnungskurs, group members only). taxPeriod is a calendar year (MWSTG Art. 34 Abs. 2) and defaults to the current one. The election carries forward until it is changed. Art. 45 Abs. 5 binds it for at least one Steuerperiode, so it may only be written while that period, and every period after it, still holds no posted foreign-currency entry. bank (Abs. 3bis) is not electable: it is the mandated fallback for currencies the ESTV publishes no rate for and stays admissible under every election.
| Name | Required | Description | Default |
|---|---|---|---|
| method | Yes | ||
| taxPeriod | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It is transparent about the election carrying forward, the minimum binding period (Art. 45 Abs. 5), the prerequisite that no posted foreign-currency entry exists, and the fact that 'bank' is a mandated fallback always admissible. These are significant non-obvious behaviors well beyond what the empty annotations or schema convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense paragraph that front-loads the core action and valid values, then adds constraints. Every sentence contributes meaningful information (defaults, carry-forward, eligibility, bank exclusion). It is somewhat legalistic with Swiss code references, which may reduce readability for a general agent, but it is concise relative to the complexity of the domain.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the legal and operational complexity, the description covers the essential facts an agent needs to call the tool correctly: parameter meanings, defaults, allowed values, main restriction, and a non-obvious fallback. It lacks any mention of return values or error behavior, but there is no output schema and this is a setter, so that is a minor gap. Overall, it is complete enough for a competent agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does for two of three parameters: 'method' options are fully enumerated with Swiss tax references, and 'taxPeriod' is explained as a calendar year with a default. The 'workspaceId' parameter is not mentioned, but it is a common multi-tenant identifier whose meaning is likely assumed from context; this slight gap prevents a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Elect the MWSTV Art. 45 conversion basis for a Steuerperiode.' It names the specific resource (the conversion basis), the verb (elect/set), and the allowed values. This unambiguously differentiates it from the sibling get_fx_method, which reads the current setting.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when the tool applies: it is used to elect a conversion method, with taxPeriod defaulting to the current year. It also specifies the condition under which it may be written (no posted foreign-currency entry in the period) and explicitly excludes the 'bank' method as not electable. However, it does not explicitly name alternative tools (e.g., get_fx_method for reading or record_exchange_rate for rates), leaving some inference work to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_opening_balancesA
Seed the workspace's opening position as ONE balanced, posted, immutable opening entry (Eröffnungsbilanz), each account carrying exactly one side in integer Rappen. Accounts are named by id or by chart number. If the set does not tie out it is REFUSED with the signed Rappen difference and nothing is written: pass differenceAccount (KMU 9100 Eröffnungsbilanz is the usual one) to book the delta explicitly, because a difference is never plugged silently. asOf defaults to the fiscal-year start. Once the prior year is closed the position is carried from that close and this verb answers carried_forward. An account that already carries a balance on that date is refused with account_already_has_balance, naming what it holds: seeding on top of it would double-count, and a bank opening balance from set_bank_opening_balance is the usual reason. A posted position has no edit path: correct it with a reversing entry (reverse_entry) and a fresh call under a new key, which this verb accepts once the storno has unwound the old position. reference records the Inventar or Beleg the position is traced to, and get_opening_balances reads it back.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| lines | Yes | ||
| reference | No | ||
| description | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| differenceAccount | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so thoroughly: immutability, no edit path, atomic refusal with nothing written, tie-out refusal with signed Rappen difference, no silent difference plugging, and explicit refusal reason account_already_has_balance. It also mentions default asOf and that reference is readable via get_opening_balances.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The opening sentence is front-loaded and every sentence contributes a distinct constraint or consequence. The definition is dense and single-paragraph rather than structured with explicit parameter mapping, but that is reasonable given the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex, irreversible write operation, the description covers success conditions, balance-refusal behavior, existing-balance refusal, carried-forward state, idempotency behavior via a new key, and the correction workflow. The read-back path is named, and no output schema exists, so the description provides enough context to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate; it explains lines as one side per account in integer Rappen, account identification by id or chart number, differenceAccount for booking the delta, asOf defaulting to fiscal-year start, and reference as Inventar/Beleg. WorkspaceId and the free-text description are left implicit, so it is strong but not fully exhaustive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb ('Seed'), a specific resource (the workspace's opening position as one balanced, posted, immutable opening entry / Eröffnungsbilanz), and the integer-Rappen constraint. It also distinguishes itself from related siblings such as set_bank_opening_balance and get_opening_balances.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly covers when to use it ('Seed the workspace's opening position'), when it cannot apply ('prior year is closed ... carried_forward', account already has a balance), and how to correct a posted position (reverse_entry then a fresh call under a new key). It also names set_bank_opening_balance as a typical source of an already-existing bank balance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_qr_auto_applyA
Switch the workspace's auto-apply dial for QR matching (P8, default OFF). ON means apply_qr_match may settle WITHOUT a per-call confirmation when, and only when, the live score is 'high' (exact reference, exact amount) for exactly the invoice being applied; every 'medium' and every override still waits for confirmed=true. This verb is deliberately not automatable (D65): a rule that could switch the unattended-money dial on would bypass the human approval the dial exists to record. A repeat call under the same idempotencyKey writes nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| autoApply | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so thoroughly. It discloses the exact behavior of ON (high-score matches settle without confirmation), that medium and overrides still wait, that the dial is intentionally not automatable, and idempotency behavior under the same key. This is exemplary transparency for a configuration tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph but packs in the essential information: purpose, conditions, non-automatability, and idempotency. It is front-loaded with the main action and then elaborates. No fluff, though it is slightly longer than strictly necessary, so 4 rather than 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple boolean switch with three parameters and no output schema, the description covers the behavioral contract, the exact conditions under which it takes effect, and idempotency. It doesn't describe return values or errors, but given the tool's simplicity and lack of output schema, this is not a significant gap. It is complete enough for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the parameters. It gives semantic context for idempotencyKey ('a repeat call writes nothing') and implies autoApply is the boolean toggling the dial, but workspaceId is never mentioned beyond being required. Since the tool name and context make autoApply and workspaceId fairly obvious, a 3 is appropriate—partial compensation but not full.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('switch') and a clear resource ('workspace's auto-apply dial for QR matching'). It also differentiates from related siblings like apply_qr_match and override_qr_match by explaining when auto-apply applies and when it doesn't. The phrase 'P8, default OFF' adds precision, making the tool's role unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the intended usage: switching the auto-apply dial, with detailed conditions for when ON actually bypasses confirmation. It also warns that the tool is deliberately not automatable (D65) to preserve human approval, which is an explicit usage constraint. However, it does not explicitly state alternatives or when NOT to use it beyond the automation warning, so it falls just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_roleB
Change a member's role. It takes effect on their next call: nothing caches a capability resolution. Demoting the only remaining owner is refused (last_owner). CONSEQUENCE: Changes what a member may see and do in the workspace; the new rights apply on their next call.
| Name | Required | Description | Default |
|---|---|---|---|
| role | Yes | ||
| memberId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing behavior. It clearly states that the role change takes effect on the next call, that no caching occurs, and that demoting the only remaining owner is refused. It also explains the consequence: changes what the member can see and do. This is strong transparency for a mutation tool, though it doesn't cover all edge cases like permissions required to invoke the tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with zero filler. The core action is front-loaded, and the two additional sentences convey critical caveats (effect timing and last_owner refusal) without redundancy. Every sentence earns its place, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (3 string parameters, no output schema), the description is largely complete. It covers the main effect, timing, and a key refusal case. It doesn't explain how to discover valid role values or what happens on invalid input, but those are likely covered by error messages or other documentation. Overall, it's sufficient for an agent to invoke the tool correctly in most scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for undocumented parameters. However, it does not explain the expected format or valid values for 'role' (e.g., role ID vs name), nor does it clarify the semantics of workspaceId or memberId beyond what is obvious from the names. The description adds minimal meaning beyond the schema, leaving an agent with ambiguity about how to supply the role parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action 'Change a member's role' with a specific verb and resource. It goes beyond a basic statement by noting the effect timing and a refusal condition, which adds clarity. However, it does not explicitly differentiate from sibling tools like define_role or revoke_member, so it doesn't fully meet the 5-criterion of distinguishing from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as define_role, revoke_member, or invite_member. It doesn't mention any prerequisites, exclusions, or conditions that would route an agent to a different tool. While it mentions a caveat about last_owner, that's a behavioral constraint, not usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_vat_methodB
Set the VAT method (effektiv/saldo) and accounting timing (soll/ist).
| Name | Required | Description | Default |
|---|---|---|---|
| vatMethod | Yes | ||
| workspaceId | Yes | ||
| vatAccounting | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It indicates a mutation ('Set') but does not state whether it requires specific permissions, whether it affects existing VAT entries, or what the response looks like. For a setter with no annotations, this is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the action and includes the key options. There is no wasted verbiage.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple setter with 3 flat parameters and no output schema, the description covers the core intent but omits details like what the values imply, any constraints, and the behavior of the workspaceId. It is adequate but minimal; an agent might not know whether the change is reversible or applies globally.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the allowed values for vatMethod and vatAccounting, adding meaning beyond the schema's bare types. However, workspaceId is not described at all, leaving that parameter without context. Partial compensation is present.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Set') and the resource (VAT method and accounting timing), and even enumerates the allowed values for each (effektiv/saldo, soll/ist). It is specific enough to understand what the tool does, though it doesn't explicitly differentiate from sibling tools that might also touch VAT settings, such as vat_configure or vat_settlement_post.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It doesn't mention prerequisites, context (e.g., before closing a period), or conditions that would make this the right choice. Sibling tools like vat_configure and vat_seed_defaults exist, but no differentiation is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_write_off_thresholdA
Set the residual below which the Studio may offer a one-click Ausbuchung, in Rappen (default CHF 1.00). It governs the offer only: a larger write-off stays recordable when it is stated deliberately. This is a product setting and not a rounding rule.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| idempotencyKey | Yes | ||
| thresholdMinor | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description takes on the full burden. It transparently discloses that the setting affects only the one-click offer, that a larger write-off remains recordable when stated deliberately, and that it is a product setting rather than a rounding rule. However, it omits details about permissions, persistence, reversibility, or side effects, which are important behavioral traits for a setter tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with no filler. The first sentence states the action, parameter, unit, and default; the second clarifies the scope; the third adds the rounding-rule caveat. Every sentence contributes unique value and the text is well front-loaded with the essential information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple setter with three parameters and no output schema, the description covers the threshold's unit, default, and behavioral scope, but leaves workspaceId semantics, idempotencyKey purpose, persistence behavior, and potential error outcomes unaddressed. It is usable but has clear gaps that an agent would need to infer or test.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain thresholdMinor's unit (Rappen) and default (CHF 1.00), providing crucial context. However, workspaceId and idempotencyKey are not described at all, leaving their roles to be inferred from naming conventions and typical patterns. This partial compensation is adequate but not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the action as setting a write-off threshold for the Studio's one-click Ausbuchung, specifies the unit (Rappen) and default (CHF 1.00), and explicitly distinguishes it from a rounding rule. This is specific and unique among siblings, leaving no ambiguity about the tool's purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage scope by stating it 'governs the offer only' and noting it is 'not a rounding rule', which gives some exclusionary guidance. However, it does not name alternative tools or provide explicit conditions for when to choose this tool over others, leaving the when-to-use guidance implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sign_requests_completeA
Markiere als signiert und lege die signierte Fassung ab: verifies originalSha256 against the requested file version (document_hash_mismatch on a swapped file), then delegates to E00 files_new_version so the signed PDF becomes a new version of the SAME file (supersedes_id, retention carried forward per OR 958f). Legal from draft (the manual wet-ink path when no provider is wired), sent and viewed (the provider path). E01 never writes file bytes itself.
| Name | Required | Description | Default |
|---|---|---|---|
| mime | No | ||
| workspaceId | Yes | ||
| signRequestId | Yes | ||
| idempotencyKey | No | ||
| originalSha256 | Yes | ||
| signedContentBase64 | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the full burden and does so thoroughly: it discloses the SHA-256 verification, the document_hash_mismatch failure on a swapped file, delegation to files_new_version, retention carry-forward per OR 958f, and the fact that E01 never writes file bytes itself. This goes well beyond what the name or schema conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the primary action, followed by the verification guard, delegation behavior, and legal states. The internal E00/E01 labels add some noise, and the single paragraph could be more scannable, but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a 6-parameter schema, no annotations, and no output schema, the description covers the critical invocation context: preconditions, main failure mode, delegation effect, state availability, and retention behavior. It lacks detail on expected return values, non-hash errors, and optional parameters, but an agent can invoke and reason about the core call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Since schema description coverage is 0%, the description must add parameter semantics, and it does for the most important ones: originalSha256 is verified against the requested file version, and signedContentBase64 is the signed PDF content that becomes a new version. Optional parameters like mime and idempotencyKey receive no explanation, but core invocation semantics are strongly clarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise action (mark as signed and store the signed version) and the exact mechanism: verify the originalSha256, then delegate to files_new_version so the signed PDF becomes a new version of the same file. This clearly separates the tool from generic file operations and sibling sign-request tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names the states where the tool is legal: draft for the manual wet-ink path with no provider wired, and sent/viewed for the provider path. It does not enumerate exclusion states or compare directly with alternatives like withdraw or delete_draft, but the context is specific enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sign_requests_createA
Fordere eine Signatur an: creates a draft sign request on an E00 file (fileId, the current head version) for one signer (signerContactId, must carry an email), at a signature level (ses or qes, the ZertES/OR Art. 14 classification), with an optional message and expiry. Writes the provider-agnostic local artifact (OP4); nothing is transmitted. One open request per signer and file (request_already_open); a superseded file version is refused (not_head_version).
| Name | Required | Description | Default |
|---|---|---|---|
| fileId | Yes | ||
| message | No | ||
| expiresAt | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No | ||
| signatureLevel | Yes | ||
| signerContactId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It reveals the side effects (writes a local artifact OP4), what it does not do (nothing transmitted), and known error conditions (request_already_open, not_head_version). This is substantial, though it omits details like idempotency behavior and return values.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one dense, front-loaded sentence that packs in purpose, constraints, and side effects without extraneous words. The mixed German/English phrasing is a minor stylistic drawback, but it does not hurt comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers most operational details: what it creates, what it does not do, and failure modes. However, there is no output schema and no description of the return value or response shape, so an agent does not know what the tool returns (e.g., a draft request ID). It also leaves idempotencyKey semantics unexplained, which matters for a create operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains the meaning of fileId (current head version), signerContactId (must have email), signatureLevel (ses or qes with ZertES/OR Art. 14 classification), and optional message/expiry. It omits workspaceId and idempotencyKey, but overall adds significant meaning beyond the bare parameter names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('creates a draft sign request') on a specific resource (an E00 file) with a single signer and signature level. It distinguishes itself from siblings like sign_requests_send by explicitly noting 'nothing is transmitted', making it unambiguous that this creates a draft only.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: it creates a draft and does not transmit anything, implying it precedes sign_requests_send. It also gives explicit constraints (one open request per signer/file, head version only). However, it does not explicitly name an alternative or say 'use sign_requests_send when ready to transmit', so it stops short of full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sign_requests_delete_draftA
Verwirf einen Signatur-Entwurf: deletes a draft request that never went anywhere. Only a draft may be deleted (invalid_transition otherwise); sent and later states are withdrawn or completed, never erased, because their event trail is evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| signRequestId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the state restriction, names the error condition invalid_transition, and explains why sent/completed requests are never erased due to their event trail being evidence. It does not mention permissions, idempotency, or what the response contains, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the action and key constraint. The opening German phrase 'Verwirf einen Signatur-Entwurf' repeats the English action and adds limited value, but the rest of the description is high-signal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple delete tool, the state-machine context is valuable and covers the most important distinction from withdraw/complete. However, with no output schema and no documented parameters, especially idempotencyKey, an agent still lacks full information about invocation details and expected results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description adds no parameter-level guidance. workspaceId and signRequestId are inferable from their names, but idempotencyKey is not explained. Because the schema is bare, the description needed to compensate for parameter meaning and did not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description specifies a clear action (delete/discard) and resource (a signature draft request), and clarifies that it applies only to drafts that 'never went anywhere'. It also distinguishes this from withdrawal/completion of sent requests, so an agent can separate it from sign_requests_withdraw and sign_requests_complete.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit condition: only a draft may be deleted, otherwise invalid_transition is returned, and sent or later states must be withdrawn or completed, not erased. This effectively routes an agent to the right action, though it does not name the sibling tools explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sign_requests_getBRead-only
Eine Signaturanfrage mit ihrem lokalen Artefakt (the OP4 envelope incl. the event trail). Reading an overdue open request persists its expiry first (lazy sweep, no background daemon in the OSS core).
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| signRequestId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already marks it as read-only, but the description adds valuable behavioral context: reading an overdue open request triggers a lazy sweep that persists its expiry, and notes there is no background daemon in the OSS core. This goes beyond what annotations provide and sets correct expectations about side effects despite the read-only hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler. The main purpose is front-loaded and the important behavioral caveat about the lazy sweep is included. The German/English mix is slightly awkward but does not waste words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 2-parameter get operation with a readOnly annotation, the description is close to sufficient. It lacks any indication of the return structure (the OP4 envelope and event trail are named but not described), and the parameter semantics are left entirely to the schema. The behavioral note is useful but the description still relies on the tool name for most of the meaning.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain what workspaceId or signRequestId mean beyond their obvious names. For a get-by-id operation these parameter names are fairly self-explanatory, but the description adds no detail about formats, required scoping, or how they relate to the returned artifact.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource ('Signaturanfrage' / sign request) and the specific verb ('get' / reading), and mentions it returns the local artifact including the OP4 envelope and event trail. It is reasonably distinct from sibling sign_requests_* tools, though the mixed German/English wording ('Signaturanfrage' with English parenthetical) adds some ambiguity about what exactly is returned.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains a specific side effect when reading an overdue open request (lazy sweep persists expiry first), which serves as an implicit warning about when to use this tool. It does not explicitly state when to choose this over sign_requests_list or sign_requests_create, nor does it provide exclusions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sign_requests_listARead-only
Die Signaturanfragen-Liste (P5): filterable by file (the per-file list in the drawer), status (the "Offene Signaturen" tracking filter) or signer. savedViewId applies a saved view (G00): its stored filters merge underneath any filter named explicitly here. Listing persists overdue expiries first (lazy sweep).
| Name | Required | Description | Default |
|---|---|---|---|
| fileId | No | ||
| status | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| signerContactId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide only readOnlyHint=trueassword, and the description adds meaningful behavioral context: it states that listing 'persists overdue expiries first (lazy sweep)' and explains how savedViewId merges filters underneath explicit filters. These are non-obvious behaviors not inferable from the schema or annotations, and they do not contradict the read-only hint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with focused details. It front-loads the core purpose and then adds necessary filter and ordering caveats. The parenthetical references (P5, G00) add a slight layer of internal jargon but do not bloat the text. Every sentence contributes information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with 5 parameters and no output schema, the description covers filtering and ordering behavior but lacks information about pagination, response format, or whether all signature requests are returned or limited. The absence of an output schema increases the need to describe return payloads onthere, which is not addressed. Still, the core filtering and saved-view semantics are well-covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description bears the burden of explaining parameters. It explains the semantics of 'file' (per-file list in the drawer), 'status' (the 'Offene Signaturen' tracking filter), 'signer', and 'savedViewId' (merging behavior). It omits 'workspaceId', but that is a common required parameter with obvious meaning, so the description compensates adequately.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states this is 'Die Signaturanfragen-Liste' (the signature request list) and enumerates filter dimensions (file, status, signer). This is a clear verb+resource statement. It does not explicitly name a sibling to differentiate from, but the function name and filter description make its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool vs. alternatives like sign_requests_get or search_global. The context about saved views and filters implies usage scenarios, but there is no explicit 'use this when...' or 'instead of...' statement. An agent would have to infer the tool's role from its name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sign_requests_record_eventA
Erfasse ein Statusereignis auf einer Signaturanfrage: viewed (sent -> viewed; recording viewed twice is a no-op), declined (sent|viewed -> declined, with optional declinedReason: also the manual "Als abgelehnt markieren" in the drawer), or expired (sent|viewed -> expired, a provider deadline event). Optional evidence is appended to the event trail of the local artifact. Every illegal edge answers invalid_transition.
| Name | Required | Description | Default |
|---|---|---|---|
| status | Yes | ||
| evidence | No | ||
| workspaceId | Yes | ||
| signRequestId | Yes | ||
| declinedReason | No | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses idempotency (viewed twice is a no-op), optional evidence appended to the event trail, and error behavior (illegal edges return invalid_transition). It also specifies allowed transitions, which is substantial behavioral disclosure, though it does not cover all potential side effects like reversibility or notifications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: purpose first, then statuses with transitions, then edge cases. Each clause adds unique information without redundancy, and the use of colons and parentheses makes it scannable. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the state machine, error responses, and evidence handling, which are the core complexities. However, it omits idempotencyKey semantics, response expectations (there is no output schema), and any prerequisite or permission details. For a 6-parameter tool with no annotations, it is fairly complete but not fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains status values and transitions, the declinedReason parameter's role, and evidence's purpose, but it does not explain idempotencyKey or the structure of the evidence object. This leaves meaningful gaps for an agent, especially since no schema descriptions exist.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool records a status event on a signature request and enumerates the specific statuses (viewed, declined, expired) with their transition semantics. It distinguishes this from siblings like sign_requests_send, sign_requests_complete, or sign_requests_withdraw by describing a distinct stateful event-recording action. The verb and resource are explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides concrete scenarios: when a request is viewed, declined (including manual marking), or expired due to provider deadline. These implied conditions give an agent clear context for when to invoke the tool, though it does not explicitly name alternatives or exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sign_requests_sendA
Sende die Signaturanfrage an den Anbieter: the ONE outbound e-sign verb (draft -> sent). P8 confirm-gated (needs_confirmation without confirmed true, for humans and agents alike), and honest about the OP4 boundary: with no e-sign provider wired (the OSS-core default) it returns needs_provider with the local artifact and the request stays a draft, for the manual wet-ink path via sign_requests_complete.
| Name | Required | Description | Default |
|---|---|---|---|
| confirmed | No | ||
| workspaceId | Yes | ||
| signRequestId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden, and it discloses key behaviors: confirm-gating (needs_confirmation unless confirmed=true), the provider-boundary failure (needs_provider with local artifact), and that the request remains a draft in that case. This is honest and actionable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action, then the caveats. Some internal jargon (P8, OP4) adds noise but the structure is efficient and every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or annotations, so the description should cover return values and edge cases; it covers the no-provider edge case but not the success return shape or idempotency semantics. The required params are in the schema, so not needed here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains the 'confirmed' parameter's effect, but workspaceId and signRequestId are only implied by names, and idempotencyKey is not mentioned at all. Partial compensation only.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('send') and resource ('signature request to the provider'), and explicitly identifies itself as 'the ONE outbound e-sign verb (draft -> sent)', distinguishing it from siblings. It is unambiguous about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains when to use: to move a draft to sent via an e-sign provider, and explicitly names the alternative sign_requests_complete for the manual wet-ink path when no provider is wired. It also notes the confirm-gating requirement, which guides the agent on setting 'confirmed'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sign_requests_withdrawA
Ziehe eine gesendete Signaturanfrage zurück: sent|viewed -> expired with expired_reason withdrawn (no seventh status is minted, §H-ENUM stays six states). A draft is not withdrawn but discarded via sign_requests_delete_draft; a terminal request answers invalid_transition.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| signRequestId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so thoroughly. It discloses the before/after state, the semantic reason, a domain invariant (no seventh status minted, §H-ENUM stays six states), and the error response for invalid transitions. This is far richer than a generic 'withdraw a request' statement.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the action, and every sentence earns its place. The state transition, enum invariant, draft routing, and terminal-error behavior are packed into two dense sentences without filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a state-machine operation with no annotations or output schema, the description covers the essential domain rules, exclusions, and error condition very well. It falls slightly short on idempotencyKey semantics and on what the caller should expect as a return value, but an agent can invoke the tool correctly with the information provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for missing parameter explanations. It does not mention workspaceId, signRequestId, or idempotencyKey, nor does it explain the role of idempotency. The domain context implies signRequestId refers to a sent request, but the parameter semantics are largely left to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Ziehe eine gesendete Signaturanfrage zurück' (withdraw a sent signature request). It then pinpoints the exact state transition (sent|viewed -> expired with reason 'withdrawn'), which clearly distinguishes this tool from sign_requests_delete_draft and other sign-request siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when the tool applies (sent|viewed requests) and when it does not: drafts must be discarded via sign_requests_delete_draft, and terminal requests yield invalid_transition. This gives an agent unambiguous selection criteria among the sign_requests_* sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stock_location_upsertA
Create or edit a stock location (Lagerort). Pass locationId to edit, omit it to create. type is an organisational tag only (warehouse/store/...), never a valuation input.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| type | No | ||
| archived | No | ||
| locationId | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It adds a valuable behavioral note that 'type' is organisational only and never a valuation input, but it does not disclose other behaviors like whether archiving is possible through archived param, idempotency semantics, or side effects on existing stock movements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no extraneous words. Front-loads the primary purpose and provides the key create/edit discriminator and a critical domain caveat. Excellent economy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter tool with no annotations and no output schema, the description is somewhat thin. It covers the create/edit switch and type caveat, but not archived behavior, workspace/idempotency usage, or expected response. Adequate for a simple upsert but leaves gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so description must compensate. It explains locationId (pass to edit, omit to create) and type (non-valuation tag), but leaves name, archived, workspaceId, and idempotencyKey semantically unaddressed. Partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Create or edit a stock location (Lagerort)' with a specific verb and resource. It also explains the create-vs-edit condition. However, it does not explicitly differentiate from sibling tools like location_create or location_update, though the 'stock' qualifier helps.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives guidance on when to pass locationId (edit) vs omit (create), but doesn't discuss when to use this tool over alternatives such as location_create/location_update or other stock tools. No explicit exclusions or alternative routing are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stock_low_stockBRead-only
The stock-tracked items whose total on-hand is at or below their D00 reorder point. Empty is not an error.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already marks this as a safe read operation. The description adds useful behavioral context beyond that: the result set is based on stock-tracked items and a specific reorder threshold, and an empty result is a valid outcome, not an error. This helps an agent avoid misinterpreting a successful empty list as a failure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The entire description is a single compact sentence. It front-loads the core selection criterion and ends with a valuable interpretive note about empty results. Every word contributes meaningful information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with one required parameter, the description covers the essential selection logic and the meaningful empty-result behavior. It does not describe the response shape, but with no output schema present, this is a minor gap. Some potential overlap with sibling low-stock tools is not addressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has only workspaceId with no description, and the description text does not mention workspaceId at all. With 0% schema description coverage, the description should compensate but offers no parameter-level meaning, leaving the agent to infer that workspaceId identifies the workspace context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies a specific resource and filter condition: stock-tracked items whose total on-hand is at or below their D00 reorder point. It reads as a noun phrase rather than an explicit imperative like 'List' or 'Retrieve,' but the intent is unambiguous. It does not differentiate from the closely named sibling inventory_low_stock, so it loses the last point.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives such as inventory_low_stock, inventory_reorder_candidates, or stock_on_hand. The statement 'Empty is not an error' hints at how to interpret results but provides no selection context or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stock_moveA
Record a stock movement (Bewegung). reason is one of receipt|issue|adjust|transfer|return: receipt/return add, issue subtracts, adjust keeps the signed qty (a stocktake shrink is negative), transfer writes a paired issue+receipt across two locations (pass toLocationId). qty is a non-zero integer. A move that would drive on-hand negative is refused with insufficient_stock unless allowNegative is set. Never posts to the ledger (OP2).
| Name | Required | Description | Default |
|---|---|---|---|
| qty | Yes | ||
| refId | No | ||
| itemId | Yes | ||
| reason | Yes | ||
| movedAt | No | ||
| refKind | No | ||
| locationId | Yes | ||
| workspaceId | Yes | ||
| toLocationId | No | ||
| allowNegative | No | ||
| unitCostMinor | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and discloses important behavior: receipt/return add, issue subtracts, adjust preserves signed qty, transfer creates a paired issue+receipt, negative stock is refused unless allowNegative is set, and it never posts to the ledger (OP2). It does not cover permissions, reversibility, or idempotency, but the stated behaviors are strong and non-obvious.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense, front-loaded sentences with no filler; every clause contributes a distinct semantic rule. The reason list and consequences are packed efficiently, making it easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex write tool with no annotations and no output schema, the description gives critical movement semantics and ledger behavior but omits idempotencyKey semantics, cost/unitCostMinor, refId/refKind, movedAt, and explicit sibling routing. It is usable for basic calls but incomplete for full correct invocation across all fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
It adds real meaning for a few key parameters: the reason enumeration, qty non-zero, toLocationId for transfers, and allowNegative's effect. However, with 12 parameters and 0% schema coverage, it leaves required fields like workspaceId, itemId, locationId, and idempotencyKey plus unitCostMinor, refId, refKind, and movedAt unaddressed, so the agent is still guessing on several fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool records a stock movement and enumerates the reason semantics (receipt/issue/adjust/transfer/return), making the core purpose evident. However, it does not explicitly differentiate itself from sibling tools like inventory_move or inventory_transfer, leaving some ambiguity about when this specific movement tool is preferred.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides no explicit when-to-use guidance or alternatives. Although the reason list implies what each action does, it never tells an agent to prefer this tool over inventory_move, inventory_transfer, or inventory_adjust, and there are no exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stock_on_handARead-only
On-hand per item x location (a P5 read model = the signed sum of stock movements), plus the location and stock-tracked-item pickers. Filter by itemId, locationId, or asOf (inclusive).
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| itemId | No | ||
| locationId | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already discloses that this is a safe read operation; the description goes beyond that by explaining that the value is a derived read model – the signed sum of stock movements – and that the asOf filter is inclusive. This adds meaningful behavioral context about the computation and filter semantics without contradicting the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loads the core function and definition, and then lists the filters. Every clause adds information and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only query with five parameters and no output schema, the description covers the essential purpose and the main filter semantics, but it leaves savedViewId unexplained and does not describe the response structure beyond mentioning that pickers are included. Given the tool's modest complexity, this is an adequate but not complete description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate for the schema's silence. It usefully explains itemId, locationId, and asOf (including that asOf is inclusive), but it leaves workspaceId implied and does not mention savedViewId at all, so two of the five parameters still lack semantic guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the tool as returning on-hand stock per item and location, and explicitly defines the read model as the signed sum of stock movements. It includes extra context about pickers and filters, but it does not explicitly name a sibling tool to differentiate it, so the agent must infer its niche among many inventory tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains that the tool can be filtered by itemId, locationId, or asOf, and notes the inclusive asOf semantics. However, it gives no explicit guidance about when to prefer this tool over related siblings such as inventory_balance_by_location, and it does not name any alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stock_run_valuationA
Run the period-end inventory valuation (Bestandesbewertung): computes the full inventory value at asOf by method (fifo|weighted_avg) and records a valuation run row with the figure and the delta against D01 own last run. REPORT-ONLY (K68): it posts NO journal entry. J06 inventory_valuation_post is the sole path inventory value reaches the books (Dr 1200 / Cr 4200); running this verb never touches the ledger, so it cannot double the Vorräte asset. A locked period is refused with period_locked before the run row is written; a re-run at the same asOf supersedes the prior run row (never edits it). Idempotent on idempotencyKey.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | Yes | ||
| method | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure—and it excels. It discloses that it never touches the ledger, cannot double the Vorräte asset, refuses locked periods with period_locked, supersedes prior run rows on re-run (never edits), and is idempotent on idempotencyKey. This is exceptional transparency beyond anything the schema could convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Although longer than average, every sentence earns its place: purpose, report-only safety, posting alternative, locked-period behavior, re-run semantics, and idempotency. The critical information is front-loaded in the first sentence, and the rest adds actionable detail without repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, absence of annotations, and absence of an output schema, the description covers all essential operational context: what it computes, what it records, side effects, failure modes, and idempotency. It even specifies the delta reference (D01 own last run). No critical gap remains for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains method values (fifo|weighted_avg), the meaning of asOf (valuation date), and the idempotency behavior tied to idempotencyKey. However, it does not explicitly describe workspaceId, leaving one of the four required parameters semantically unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Run'), a precise resource ('period-end inventory valuation'), and the action in detail: computes inventory value at asOf by method (fifo|weighted_avg) and records a valuation run row with the figure and delta. It clearly distinguishes itself from siblings like inventory_valuation_post and stock_valuation_report by emphasizing it is report-only and posts no journal entry.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use and when-not-to-use guidance: it is for report-only valuation, and it explicitly names inventory_valuation_post as the sole posting path to the ledger. It warns about locked periods and re-run semantics, leaving no ambiguity about when or how it should be used.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stock_stocktake_commitA
Commit an Inventur: posts every difference in one batch, each minting a stock movement through the same movement path (reason:adjust, moved_at=frozenAt, ref_kind:stocktake), and seals the session as the committed Bestandesnachweis (OR 958c Abs. 2). No ledger reach: any financial effect arrives later through stock_run_valuation. Refused with uncounted_lines if any line is not counted; idempotent, differences never post twice.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and goes well beyond the name: it discloses movement-path attributes, batch semantics, the sealing action, the later valuation dependency, the uncounted_lines refusal, and idempotency. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with no filler: purpose and mechanism, ledger implication, then failure mode and idempotency. Every clause adds non-redundant information and the core action is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutating commit operation with no output schema, the description covers batch mechanics, movement reason/timestamp/ref_kind, sealing outcome, no-direct-ledger behavior, a failure condition, and retry safety. An agent has enough to decide when and how to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It indirectly explains sessionId through 'session' and idempotencyKey through 'idempotent, differences never post twice', but workspaceId is never described and the parameters are not explicitly mapped to their roles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: 'Commit an Inventur' and the stocktake session, then specifies commit-specific behavior: posting every difference in one batch and sealing the session. It distinguishes itself from stocktake open/count/report siblings by describing the sealing and batch execution details.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear preconditions and safe usage context: refusal if any line is uncounted, idempotency for retries, and no ledger reach with financial effects routed through stock_run_valuation later. It does not explicitly say 'use this when the stocktake is fully counted' or compare against the inventory_stocktake_commit sibling, so it falls just short of fully explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stock_stocktake_countA
Record the counted qty on one stocktake line (item x location). Only while the session is open. Setting counted to the book qty is how a line is confirmed unchanged.
| Name | Required | Description | Default |
|---|---|---|---|
| itemId | Yes | ||
| sessionId | Yes | ||
| countedQty | Yes | ||
| locationId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full behavioral burden. It discloses the mutation and the open-session constraint, and explains the book-qty confirmation semantics. However, it omits other details like idempotency, whether existing counts are overwritten, permissions required, or error behavior beyond an invalid session state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two terse sentences, no filler. The core action is front-loaded, followed by a critical usage constraint and a semantic hint about book qty. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter, no-output-schema tool with no annotations, the description provides the essential purpose and a key rule, but leaves gaps: it does not explain the expected return value, whether the call is idempotent, how to obtain valid identifiers, or what happens if the session is closed mid-call. It is adequate for a simple mutation but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must substitute. It explains that 'counted qty' is the value being set and clarifies 'item x location' as the line identity, which maps to itemId and locationId. But workspaceId is not mentioned, and the relationship among sessionId, itemId, and locationId as a composite key is only implied, not explicitly specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action: 'Record the counted qty on one stocktake line (item x location).' It uses a precise verb and resource, and distinguishes this from whole-stocktake operations by emphasizing 'one line.' The phrase 'item x location' ties directly to the itemId and locationId parameters.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit timing condition ('Only while the session is open') and a key usage rule ('Setting counted to the book qty is how a line is confirmed unchanged'). It does not explicitly name alternative sibling tools or state when not to use it, but the context is clear for this specific mutation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stock_stocktake_openA
Open an Inventur (physical stocktake): snapshots book qty per item x location as of frozenAt (the balance-sheet date) into a session, optionally scoped to one location. The frozen book qty is the on-hand read model captured at that instant.
| Name | Required | Description | Default |
|---|---|---|---|
| frozenAt | Yes | ||
| locationId | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears full responsibility for behavioral disclosure. It discloses that the tool snapshots book quantities into a session and that the frozen quantity is the on-hand read model at that instant, giving a clear picture of the data captured. However, it does not mention side effects such as reversibility or locking, but the snapshot and freeze concepts are explicitly described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no extraneous detail, and the core action is front-loaded. It earns a high score for efficiency and clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema or annotations, and the schema descriptions are absent. The description lacks key information such as what the tool returns (e.g., session ID), how to use the session in subsequent steps, and explicit guidance on the required parameters beyond two. This makes it incomplete for an agent to correctly invoke the tool in a multi-step workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the meaning of frozenAt ('balance-sheet date') and locationId ('optionally scoped to one location'), but it does not explain workspaceId or idempotencyKey, both of which are required. This partial coverage is helpful but leaves a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('open') and resource ('Inventur/physical stocktake'), and explains the core action of snapshotting book quantities per item/location at a frozen date. This distinguishes it from other stocktake tools in the sibling list, such as inventory_stocktake_count or inventory_stocktake_commit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the usage context (opening a physical stocktake session) and mentions optional location scoping, but it does not explicitly state when to use this tool versus alternatives like stock_stocktake_create or inventory_stocktake_create. No exclusions or alternative references are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stock_stocktake_reportCRead-only
The stocktake diff read model (P5): book vs counted per line with the computed difference, and the uncounted/over/under tallies.
| Name | Required | Description | Default |
|---|---|---|---|
| sessionId | Yes | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already covers the read-only nature, so the description does not need to restate that. It adds value by describing the output content (line-level differences and tallies). However, it does not disclose behavioral prerequisites like whether a session must exist or be in a certain state, nor any pagination or filtering behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no fluff; it front-loads the key concept ('stocktake diff read model') and packs useful detail about the data. The 'P5' parenthetical is cryptic and unexplained, but overall the structure is efficient and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of output schema, zero parameter descriptions, and no usage guidance, this description is not sufficient for an agent to confidently call the tool. It does not clarify what a valid session is, what the response shape looks like, or how this report fits into the stocktake workflow. The tool name and readOnly annotation provide some context, but the description leaves too much to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention any parameter. It does not explain what workspaceId, sessionId, or savedViewId refer to in the context of a stocktake report. The agent must guess that sessionId likely identifies a stocktake session, but no explicit semantic mapping is provided, so the description fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource ('stocktake diff read model') and indicates the core content: book vs counted per line with computed differences and uncounted/over/under tallies. It is distinguishable from sibling stocktake tools like stock_stocktake_open or stock_stocktake_commit. However, 'P5' is unexplained jargon and there is no explicit verb (get/report), making the purpose slightly less crisp.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives. The description does not state prerequisites (e.g., an open or counted stocktake session) or exclude scenarios where other stocktake report tools would be more appropriate. The agent is left to infer usage context from the tool name and siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stock_valuation_reportARead-only
The valuation read model (P5): per-item quantity, unit cost and value at asOf by method, the value already posted to the ledger and the unposted delta, the OR 960c lower-of-cost-or-market flag, a Stetigkeit warning when the method differs from the last run, and the linked committed Inventur (OR 958c Abs. 2 Bestandesnachweis) when one exists at asOf.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | Yes | ||
| method | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, and the description is consistent. It adds behavioral context beyond the annotation: it is a read model, it computes posted vs unposted delta, it emits a Stetigkeit warning conditionally, and it links a committed Inventur when one exists. This tells the agent the call is side-effect-free and what conditional outputs to expect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence with no filler. It front-loads the core report contents and then appends the regulatory and warning flags. It is efficient, though the long list of clauses makes it slightly harder to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the report tool has no output schema, the description thoroughly enumerates the report sections and conditional elements. The main gap is parameter semantics: method values, asOf format, and any default behavior are not specified. Still, for a read-only report with rich output, it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden for parameter meaning. It clarifies that method determines valuation method and asOf is the valuation date, but it does not specify allowed method values or date format. It adds partial meaning beyond the raw schema, but not enough for a parameter-complete understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it is the valuation read model (P5) for stock. It enumerates exact outputs (quantity, unit cost, value, posted/unposted delta, OR 960c flag, Stetigkeit warning, linked Inventur), which makes it distinguishable from sibling report tools like inventory_valuation_report or stock_run_valuation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: call this tool when you need a valuation report at a given asOf date by method, including LCM and consistency warnings. However, it does not explicitly state when not to use it or point to alternatives, leaving the selection among stock_valuation_report, stock_run_valuation, and inventory_valuation_report to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_matchesARead-only
Propose a match for every txn of an imported statement, writing nothing. A routed credit shows A21's live score (an invoice, or none). A debit or a reversed entry is RANKED (A36): the exact amount + currency gate is mandatory, then value-date proximity (workspace window, default +/- 5 days), counterparty-name overlap and reference hits order the gated bills, each proposal carrying a signals[] list (amount/value_date/counterparty/reference) and its confidence. A debit whose batch PmtInfId matches a generated A18 payment batch proposes the batch itself (kind:payment_batch); a batch-total mismatch is shown blocked (one-click disabled, use the manual split). Each txn also carries needsReview: true when no proposal reaches the workspace review threshold.
| Name | Required | Description | Default |
|---|---|---|---|
| statementId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite the readOnlyHint annotation already covering safety, this description goes far beyond with algorithm details: mandatory amount+currency gate, value-date window, ordering criteria, signals[] composition, confidence, batch matching (A18), and the needsReview flag. It even discloses edge cases like batch-total mismatch and blocked proposals. No contradictions with the readOnlyHint are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but structured logically: starting with the overall purpose, then credit handling, debit ranking, batch special cases, and a closing note on needsReview. It is lengthy due to the complex behavior it must communicate, but each clause adds value and it front-loads the primary action. It could be broken into bullets, but remains appropriately sized for a complex matching engine.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description takes care to enumerate important return elements: signals[] fields, confidence, kind:payment_batch, blocked proposals, and needsReview. It covers the main decision logic and edge cases. Minor gaps exist around exact response shape or error conditions, but the essential information an agent needs to interpret results is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has only workspaceId and statementId with zero description coverage ×. The description indirectly references these ('workspace window', 'imported statement') but never explicitly maps them to parameters. Parameter names are self-explanatory, yet the description doesn't compensate fully for the missing schema descriptions – it provides context but not direct definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Propose a match for every txn of an imported statement'. It clearly distinguishes this from write operations by stating 'writing nothing', and the detailed behavior (A21, A36, payment_batch) makes the purpose unmistakable. It also differentiates from siblings like confirm_match or review_bank_txn by placing the emphasis on suggestion rather than confirmation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly establishes the context for use: matching transactions from an imported statement after import, while 'writing nothing' indicates a read-only pre-step. It doesn't explicitly name alternatives like suggest_payment_matches or state when not to use it, but the focus on imported statements vs. other matching flows gives a clear contextual trigger.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_payment_matchesBRead-only
Rank the open items a payment might settle, with the reason in words: reference matches, amount and customer match, or amount close. A QR or Creditor Reference whose check digit fails is reported as a typo and never ranked as a match, and a candidate that fits no tier carries no reason at all. This is the same read model the Studio's candidate list renders.
| Name | Required | Description | Default |
|---|---|---|---|
| currency | No | ||
| direction | No | ||
| reference | No | ||
| amountMinor | No | ||
| workspaceId | Yes | ||
| counterpartyId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds meaningful behavioral details beyond the readOnlyHint annotation: it explains that check-digit failures are reported as typos and never ranked, and that candidates fitting no tier carry no reason. It also notes the tool uses the same read model as the Studio's candidate list, giving agents insight into expected behavior. However, it doesn't clarify what the tiers are or how the ranking order works.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose, and packs substantial detail about behavioral nuances. It is efficient and avoids fluff, though the second sentence is dense and could be clearer about tier definitions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 6 parameters, no schema descriptions, and no output schema, the description is incomplete. It does not describe the output structure, required fields, or the meaning of parameters like direction or amountMinor. It also leaves the notion of 'tiers' undefined. An agent would struggle to use this tool correctly without additional documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not explain any of the six parameters (currency, direction, reference, amountMinor, workspaceId, counterpartyId). With 0% schema description coverage, the description should compensate but only loosely mentions 'reference' and 'amount' without mapping them to parameters. This is a significant gap for an agent trying to construct valid calls.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool ranks open items a payment might settle, with reasons in words. It mentions reference matches, amount and customer match, or amount close, and notes check digit failures are treated as typos. However, it does not explicitly differentiate from the sibling 'suggest_matches', relying on the phrase 'same read model' which is somewhat indirect.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no explicit guidance on when to use this tool versus alternatives like suggest_matches, preview_payment, or match_qr_payment. It implies a read-only suggestion role but does not state when it is the right choice or what distinguishes it from other matching/suggestion tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
supplier_performance_alertsARead-only
Offene Leistungs-Warnungen (performance alerts, US-I05.5): for every supplier active in the window, each configured alert threshold a metric currently breaches (by default overall score below 70 or OTIF below 85) becomes an open alert with the supplier, metric, current value, threshold and period. Alerts are DERIVED and recomputed on every read, not stored as tickets. Pass onlyOpen:false to also see the metrics that were evaluated and did not breach. Posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| onlyOpen | No | ||
| windowDays | No | ||
| workspaceId | Yes | ||
| configOverride | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description states that alerts are DERIVED, recomputed on every read, and not persisted as tickets, and explicitly says 'Posts nothing.' This meaningfully clarifies side effects and data freshness without contradicting the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences cover purpose, behavior, and a key parameter. The German label and requirement ID add a little noise, but the content is front-loaded and every sentence contributes information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description names the output fields (supplier, metric, current value, threshold, period) and explains derivation, which is valuable given the lack of an output schema. It misses details on configOverride and asOf semantics, but the core call shape and behavior are sufficiently clear for a read-only listing tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must carry parameter semantics and does explain onlyOpen behavior and the window concept, and references configurable thresholds. However, it does not document asOf, windowDays units/defaults, or the shape of configOverride, leaving significant gaps for a five-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as returning open supplier performance alerts for threshold breaches, with a concrete resource (suppliers/metrics) and output fields. It distinguishes itself from sibling analytics tools by emphasizing that alerts are derived and not stored as tickets.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear scope (suppliers active in window, configured thresholds) and parameter guidance (onlyOpen:false for non-breached evaluations), but does not explicitly say when to prefer this over sibling supplier-performance tools such as supplier_performance_rank or supplier_performance_trend. The usage context is implied rather than stated as a routing rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
supplier_performance_explainARead-only
Die Herleitung einer Kennzahl (metric explanation, US-I05.3): the exact formula applied, the workspace tolerance / weight values used, and the full list of contributing source documents (goods-receipt ids and po-line ids for the delivery metrics, po_match ids for price and override) with their per-row values. This is the audit view a Treuhänder reads to reconstruct any number on the scorecard down to the Beleg. Re-running it with the same live data yields the identical result.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| metric | Yes | ||
| supplierId | Yes | ||
| windowDays | No | ||
| workspaceId | Yes | ||
| configOverride | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint: true, and the description adds behavioral details beyond that: it states that re-running with the same live data yields identical results (determinism), and it describes the audit scope (full list of contributing documents). It does not contradict annotations and provides useful behavioral context about the tool's reliability and output scope.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise (two sentences) and front-loaded with the core purpose. It includes relevant details about output content without excessive fluff. The use of German in the first phrase adds a slight redundancy with the English explanation, but overall it is well-structured and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 7 parameters (3 required) and no output schema, the description is incomplete for correct invocation. It explains what the tool returns (formula, tolerances, source documents) but does not explain how parameters affect the output, what the optional parameters (to, from, windowDays, configOverride) mean, or the structure of the response. An agent would need additional context to call it correctly with appropriate parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining parameters. It does not describe any of the 7 parameters (workspaceId, supplierId, metric, to, from, windowDays, configOverride). Even the required parameters are only implied by the context (metric, supplier, workspace) but not explicitly documented. The description focuses entirely on the return value and fails to provide parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: it explains the derivation of a metric by showing the exact formula, tolerance/weight values, and contributing source documents with per-row values. This is a specific verb+resource (explain supplier performance metric) and clearly distinguishes it from sibling tools like supplier_scorecard_get or supplier_performance_rank, which retrieve scorecard data rather than audit derivations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: it is the audit view a Treuhänder uses to reconstruct any scorecard number down to the source documents. It implies when to use it (when an audit trail or detailed derivation is needed) and mentions deterministic re-runs, but it does not explicitly name alternative tools or conditions for not using it. However, the context is strong enough for an agent to infer appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
supplier_performance_rankARead-only
Rangliste der Lieferanten (supplier ranking) by one metric over a window (default the overall score): each row carries the supplier, the metric value, the overall score, the activity count and the delta versus the previous window. Suppliers below minActivity receipts are moved to a separate insufficient[] list rather than ranked on thin data. Ordering is stable (ties break by supplier name then id) and defaults to best-first for the chosen metric. Pure derivation over live receipts and matches; posts nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| limit | No | ||
| order | No | ||
| metric | No | ||
| windowDays | No | ||
| minActivity | No | ||
| workspaceId | Yes | ||
| configOverride | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide readOnlyHint=true, and the description reinforces this with 'posts nothing' and 'pure derivation'. It adds meaningful behavioral detail beyond annotations: the insufficient[] list separation, stable ordering with tie-break rules, default best-first ordering, and the fact that it derives from live receipts and matches. This is strong transparency for a read-only reporting tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: it states the output shape, the threshold behavior, ordering rules, and side-effect profile in four sentences. Every sentence adds distinct information and the most important facts (ranking, window, default metric) are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only ranking tool with no output schema, the description covers the key output fields, ordering, threshold behavior, and side-effect safety. It does not document the exact format of date parameters (from/to) or configOverride, but the core invocation context is sufficiently clear for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden. It explains the meaning of metric, window, minActivity, and order implicitly (default best-first, stable ties), but does not explain to, from, limit, configOverride, or workspaceId. The description adds value for the core ranking parameters but leaves several parameters undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool produces a supplier ranking (Rangliste der Lieferanten) by one metric over a window, with each row carrying specific fields. It distinguishes itself from sibling tools like supplier_performance_trend and supplier_scorecard_get by specifying the ranking/derivation nature and the separate insufficient[] list behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the default metric (overall score), the window concept, and the minActivity threshold behavior, which tells an agent when this tool is appropriate for ranking suppliers. It does not explicitly name alternative tools or state when NOT to use it, but the context of 'pure derivation over live receipts and matches; posts nothing' plus the sibling names gives reasonable usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
supplier_performance_trendARead-only
Der Trend einer Kennzahl (metric trend) for one supplier over the last N consecutive windows, oldest first: each point is a pure re-evaluation of the chosen metric for that window, and a window with no activity reports value null (a genuine gap), never a fabricated zero. Use it for the scorecard sparkline or to see whether delivery reliability is improving. Derived, not stored.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| metric | Yes | ||
| periods | No | ||
| supplierId | Yes | ||
| windowDays | No | ||
| workspaceId | Yes | ||
| configOverride | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description adds valuable behavioral details: each point is a pure re-evaluation, missing windows are genuine nulls rather than fabricated zeros, and the data is derived rather than stored. This richness helps the agent anticipate edge cases and trust the output semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core purpose. The null-gap behavior and derived-not-stored distinction add high-value information without padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The core purpose and key output semantics are clear, but the tool has 7 parameters including a nested configOverride and no output schema. The description does not explain how periods and windowDays interact, what configOverride does, or the response shape, so an agent is left guessing on several invocation details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description only obliquely references supplier, metric, and windows. It does not explain periods, windowDays, to, or configOverride, leaving 7 parameters largely underspecified for a tool with no property descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly defines the tool as computing a metric trend for one supplier over consecutive windows, with explicit direction (oldest first). It distinguishes itself well from closely related siblings like supplier_performance_rank or supplier_scorecard_get by focusing on trend visualization and improvement assessment.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It names concrete use cases: scorecard sparkline and checking whether delivery reliability is improving. However, it doesn't explicitly say when to prefer an alternative tool, so while usage context is clear, exclusions are absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
supplier_price_listARead-only
Liste die Lieferantenpreise (P5): the append-only price history for a supplier/item filter. When BOTH a supplier and an item are named, the ONE effective resolved price at at (including the item-cost fallback with source:item_cost) rides the read, so an agent can price a reorder in one call.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | ||
| itemId | No | ||
| workspaceId | Yes | ||
| supplierContactId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description adds meaningful behavior beyond that: the price history is append-only, and both-filter queries return one effective resolved price including an item-cost fallback with source:item_cost. This is useful context. The phrase 'rides the read' is jargon, but the overall behavioral disclosure is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact: two sentences, with the core purpose stated first and the special resolved-price behavior front-loaded in the second sentence. Every sentence contributes. Minor jargon like 'rides the read' and the unexplained 'P5' reduce clarity slightly, but the structure is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and 4 parameters, the description should be more explicit about return shape and edge cases. It explains the both-filters-effective-price mode well but does not describe what a normal history response looks like, what happens with only one filter, or the meaning of workspaceId. It is adequate for a basic call but incomplete for robust agent use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of explaining parameters. It does explain the supplier/item filter and references `at` as the effective-price timestamp, and mentions the item-cost fallback. However, it does not define each parameter explicitly, nor explain what happens when only one filter is provided or what format `at` should take.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource (supplier prices / P5 history) and a concrete behavior: listing append-only price history with a supplier/item filter. It also describes the special one-effective-price mode when both filters are present, which distinguishes it from a plain list. However, it does not explicitly contrast itself with related siblings like price_resolve or supplier_price_upsert, and the 'P5' parenthetical is unexplained.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear conditional rule: when both supplier and item are named, the tool returns the single effective resolved price at `at`, so it can be used to price a reorder. This effectively tells an agent when the resolved-price mode is appropriate. It does not explicitly state when NOT to use this tool or name alternatives, so it stops short of a full usage guideline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
supplier_price_upsertA
Erfasse einen Lieferantenpreis (append-only Historie nach valid_from): persists a supplier_item_price row (supplierContactId -> C00, itemId -> D00, supplierSku?, priceRappen, currency, validFrom, leadTimeDays?), mirroring D00s price-list discipline. po_upsert pre-fills each lines price (and derives expected_on from the longest lead time) from resolveSupplierPrice. A supplier that is not a C00 contact or an unknown itemId is refused (invalid_reference); priceRappen < 0 is refused.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| itemId | Yes | ||
| currency | No | ||
| validFrom | Yes | ||
| priceRappen | Yes | ||
| supplierSku | No | ||
| workspaceId | Yes | ||
| leadTimeDays | No | ||
| idempotencyKey | No | ||
| supplierContactId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral disclosure. It reveals append-only history semantics, refusal of invalid references, and rejection of negative prices. It also notes the relationship to D00's price-list discipline. This is substantial behavioral context, though it omits details like idempotency behavior or return values.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description packs essential information but in a dense run-on sentence that mixes German and English. It is not elegantly structured and could be split into clearer segments. Some redundancy exists (e.g., German phrase followed by English explanation), but the content is relevant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 10 parameters, no output schema, and no annotations. The description covers core domain rules and validation errors, but it does not describe the return value, behaviors around idempotencyKey, actor, or currency defaults. It is adequate for basic invocation but leaves some operational questions unanswered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains supplierContactId must be a C00 contact, itemId must be a D00 item, supplierSku and leadTimeDays are optional, validFrom drives append-only history, and priceRappen must be non-negative. These mappings add meaning beyond the bare schema, though parameters like actor, workspaceId, and idempotencyKey are left unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool records a supplier price and persists a supplier_item_price row. The verb and resource are specific, and the append-only and mirroring details add precision. However, it does not explicitly distinguish itself from related sibling tools like price_lists_upsert or supplier_price_list, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used to record a supplier price and mentions po_upsert's reliance on resolveSupplierPrice for pre-filling prices. This gives context, but it does not explicitly state when to use this tool versus alternatives or when not to use it. No exclusions are given beyond validation errors.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
supplier_scorecard_getARead-only
Die Lieferanten-Scorecard (supplier scorecard) for one supplier and window: OTIF, on-time %, quantity fill %, average delivery delay, quantity variance %, price variance %, match-override rate, inspection-rejection rate and a weight-normalised overall score 0-100, each with a green/amber/red light, plus the previous-window comparison and the top contributing exceptions (late deliveries, short shipments, price overrides) with deep-link PO / receipt / match ids. Every number is DERIVED from the live I02 goods receipts and the D02 three-way-match rows and is reproducible (same data, same result); nothing is posted. Empty period answers ok with empty:true, never an error. A supplier id from another workspace is not_found (tenant isolation). Pass configOverride for a what-if with different weights or tolerances.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| supplierId | Yes | ||
| windowDays | No | ||
| workspaceId | Yes | ||
| configOverride | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes far beyond the readOnlyHint annotation: states nothing is posted, data is derived and reproducible, empty periods return empty:true, tenant isolation yields not_found, and configOverride enables what-if analysis. These are non-obvious behavioral details an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence that front-loads the core purpose and then details output and edge cases. Every clause adds relevant information, though it could be broken into clearer segments for readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers key output elements (metrics, lights, comparison, exceptions with deep-links) and edge cases, but lacks explicit return format structure or parameter syntax, which is critical given no output schema and 6 parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema has 0% description coverage, and the description only vaguely references 'window' and 'configOverride' without explaining format, units, or required relationships. It fails to compensate for the lack of parameter documentation, especially for to/from/windowDays.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns a supplier scorecard for one supplier and time window, listing specific metrics and comparisons. It unambiguously identifies the resource and scope, distinguishing it from broader procurement tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides context (single supplier, window) and mentions configOverride for what-if scenarios, but does not explicitly differentiate from sibling tools like supplier_performance_trend or supplier_performance_rank. No clear when-not-to-use guidance or alternative selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sync_artifact_readARead-only
Resolve a published artifact.* handle by its content hash (sha256), returning the stable reference the blob is addressed by. The stream carries handles, never inline bytes; the blob itself is fetched from the content store the runtime binds. A hash that is not a published handle in this workspace is artifact_not_found. Requires the sync.read capability.
| Name | Required | Description | Default |
|---|---|---|---|
| sha256 | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true, so the description correctly does not repeat that. It adds valuable behavioral context: the stream carries handles, not inline bytes, and the blob is fetched from a bound content store. It also discloses the error behavior (artifact_not_found) and the required capability, which goes beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose and then adding behavioral and error context. Every sentence contributes meaning without waste, and the structure is clear and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a resolver with no output schema, the description covers the essential aspects: what it does, how it behaves (returns handles), error conditions, capability requirements, and scope. It does not explicitly describe the return type format, but it is sufficient for an agent to invoke correctly. Given the annotations and simple parameter set, this is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It explains that sha256 is a content hash and that the workspace scope is implied by 'in this workspace'. However, it does not explicitly define workspaceId, and it lacks detailed parameter descriptions. The description adds some meaning but does not fully compensate for the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool resolves a published artifact handle by its content hash (sha256) and returns the stable reference. The verb 'resolve' and the resource 'artifact handle' are specific, and it distinguishes from siblings like sync_stream_read and sync_publish_enable by focusing on resolution rather than streaming or publishing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides useful context: it clarifies the tool returns handles, not inline bytes, and that the blob is fetched from the content store. It also states the required capability and the error condition for missing hashes. However, it does not explicitly mention alternatives or when to use this over sibling tools like sync_stream_read, leaving usage guidance implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sync_publish_disableA
Turn the till-sync publish stream OFF for this workspace. Stops new appends and refuses stream reads (publishing_disabled), but the already-published outbox rows remain, append-only, for audit; the epoch is kept so re-enabling resumes the same stream. The stop button.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so excellently. It discloses the exact consequences: publishing is disabled, new appends stop, stream reads are refused, already-published outbox rows remain append-only for audit, and the epoch is preserved for later re-enabling. This far exceeds minimal mutation disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the main purpose appears in the first sentence, followed by important behavioral details and a memorable closing ('The stop button'). Every sentence earns its place with no fluff or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter mutation with no output schema and no annotations, the description covers the core behavioral contract well. The only real completeness gap is the lack of idempotencyKey semantics, which prevents it from being fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for both required parameters. It only addresses workspaceId implicitly via 'for this workspace', and says nothing about idempotencyKey, its purpose, or how it should be generated/reused. This is a meaningful gap for an agent preparing the call.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action ('Turn the till-sync publish stream OFF') on a specific resource ('for this workspace'), and details the exact behavior: stops new appends and refuses stream reads. It also implicitly distinguishes itself from related sibling tools like sync_publish_enable and sync_stream_read by describing the off-state semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides strong contextual guidance: use this when you need to stop the publish stream for a workspace, and that re-enabling will resume the same stream. It does not explicitly name alternatives or exclusion conditions, so it falls short of a full 5, but the context is clear enough for an agent to select it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sync_publish_enableA
Turn the till-sync publish stream ON for this workspace (owner dial, default OFF: a local install publishes nothing). Enabling is the consent act for any egress; it mints the stream epoch once and backfills existing posted history so the stream is complete. Reversible with sync_publish_disable.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It discloses that enabling is a consent act, mints the stream epoch once, and backfills history, which are important side effects. It also notes the default OFF state and reversibility. It does not detail potential errors or the exact impact of repeated calls, but the idempotencyKey hints at idempotency, and the core behaviors are well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. It front-loads the primary action and then efficiently covers the consent act, epoch minting, backfill, and reversibility. Every sentence adds value, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the purpose, main behavior, and reversibility, but it omits parameter explanations (especially idempotencyKey) and does not describe the return value or error conditions. Given the tool's simplicity, it is largely adequate, but the missing parameter semantics and lack of output expectations leave gaps that could affect correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must explain parameters. It only mentions 'this workspace' which loosely maps to workspaceId, but it does not explain idempotencyKey at all. No syntax, purpose, or format is provided beyond the parameter names, leaving the agent to guess at the meaning and constraints of idempotencyKey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action: turning on the till-sync publish stream for a specific workspace. It specifies the resource (till-sync publish stream), the action (turn ON), and the scope (this workspace), and distinguishes it from its disable counterpart. It also explains the default state and the consent aspect, making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it: to enable publishing, noting that it is the consent act for egress. It explicitly mentions reversibility via sync_publish_disable, which frames the alternative. However, it does not state explicit preconditions or when not to use it, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sync_stream_readARead-only
Read the append-only publish stream from a consumer-held cursor {seq, epoch}, up to limit events, with the current headSeq so you know your lag. At-least-once: a re-read from the same cursor is always safe. A cursor beyond the head is invalid_cursor; a cursor from a superseded epoch (after a restore) is cursor_reset_required. Requires publishing to be on and the sync.read capability.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| cursor | No | ||
| workspaceId | Yes | ||
| contractVersion | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, so the read-only nature is already known. The description adds valuable behavioral context: at-least-once guarantee, specific error conditions (invalid_cursor, cursor_reset_required), and prerequisites (publishing enabled, capability). This goes well beyond annotations and provides the agent with important operational details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tightly packed sentences with no fluff. The main purpose and key behavior are front-loaded; every sentence adds distinct information about semantics, safety, or errors.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, but the description mentions returning headSeq for lag awareness skipping the full response format. It covers error states, requirements, and semantics. Missing details on contractVersion and the exact nature of limit (required vs optional, upper bound) keep it slightly below complete, but it is still strong given the context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies 'cursor' as {seq, epoch} and 'limit' as 'up to limit events'. However, it omits the meaning of contractVersion and does not explicitly document workspaceId beyond its requirement. Partial compensation for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('read'), a clear resource ('append-only publish stream'), and precise mechanics: consumer-held cursor {seq, epoch}, limit events, and returning headSeq for lag. This distinguishes it from siblings like sync_stream_status (status) and sync_artifact_read (artifact-specific read).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: reading with a cursor, at-least-once semantics making re-reads safe, and explicit requirements ('publishing to be on and the sync.read capability'). It does not explicitly name alternative tools or exclusions, but the usage context is sufficiently conveyed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sync_stream_statusARead-only
The monitoring read: whether publishing is on, the stream head and epoch, and (when you pass your cursor) the lag = head - cursor.seq. Operational metadata only, no tenant rows. Requires the sync.read capability.
| Name | Required | Description | Default |
|---|---|---|---|
| cursor | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already flags this as a safe read; the description adds substantive behavior on top: it returns operational metadata only (publishing flag, stream head, epoch), computes lag only when a cursor is passed, and explicitly excludes tenant rows. The sync.read capability requirement and the lag formula are useful context beyond what the annotation provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the first front-loads the purpose and outputs, the second narrows scope with an exclusion, and the third states the required capability. No filler and no repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter read tool with a readOnly annotation, the description is nearly complete: it explains what will be returned, the triggering condition for the optional argument, the exclusions, and the auth requirement. The only gaps are the unexplained `workspaceId` and the unstated meaning of 'epoch,' neither of which is fatal for a monitoring read.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter semantics. It gives real meaning to `cursor`: optional, and its presence triggers the lag computation (lag = head - cursor.seq). But `workspaceId`, the required parameter, is never mentioned in the description, so its purpose and constraints remain unexplained despite the schema only labeling it as a string.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('monitoring read') and resource (stream status), and enumerates exactly what it reports: publishing on/off, stream head, epoch, and lag computed as head - cursor.seq. It also differentiates itself by scope ('operational metadata only, no tenant rows'), distinguishing it from sibling sync tools that deal with tenant data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes clear context: this is the read/monitor variant within the sync_* family, distinct from sync_publish_enable/sync_publish_disable which mutate state. The exclusion 'no tenant rows' implies a sibling (sync_stream_read) handles actual data reads, and 'Requires the sync.read capability' states the prerequisite. However, no alternative is named explicitly, so routing guidance is inferred rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_cancelA
Brich eine Aufgabe ab: the terminal "wird nicht erledigt", distinct from done. No activity is logged and no recurrence occurrence is spawned; a cancelled task returns only through an explicit status filter.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly states that cancellation is terminal, that no activity is logged, that no recurrence occurrence is spawned, and that cancelled tasks require an explicit status filter to be seen. These are important behavioral traits beyond the basic 'cancel' action, making the tool's side effects transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded: it opens with the core action, then adds essential behavioral distinctions and consequences. Every sentence adds value with no fluff, making it an efficient and well-structured description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple cancel operation with no output schema, the description covers key behavioral aspects well, but it is incomplete due to the lack of parameter explanations (especially idempotencyKey) and does not address error conditions or the response format. Given the schema coverage is 0%, the description should compensate more for parameters, but it does not, leaving gaps in the full context an agent needs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, so the description must compensate for parameter meaning. The description does not mention any of the parameters (taskId, workspaceId, idempotencyKey) at all. While taskId and workspaceId are self-explanatory from their names, idempotencyKey is non-obvious and lacks any explanation, leaving a significant gap in parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (cancel a task) and the resource (task), and explicitly distinguishes it from 'done', which differentiates it from the sibling tasks_complete tool. The verb and scope are specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool versus alternatives by stating it is distinct from 'done' and by describing the behavioral consequence (cancelled tasks only appear via explicit status filter). However, it does not explicitly name the alternative tool or provide a direct comparison, so it falls short of fully explicit usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_completeA
Erledige eine Aufgabe: sets done and stamps completedAt; with logActivity true on a contact-linked task, appends one OP5 timeline entry through contacts_log_activity; on a recurring task, mints the next occurrence anchored on the completed due date (never on completion time) and answers spawnedTaskId, or seriesEnded when UNTIL or COUNT closes the series. One idempotency key covers complete, log and spawn together. Allowed for holders of tasks.write and for the task assignee.
| Name | Required | Description | Default |
|---|---|---|---|
| taskId | Yes | ||
| logActivity | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so thoroughly. It discloses the side effects (timeline entry via contacts_log_activity, next occurrence minting), the anchoring rule for recurring tasks (completed due date, never completion time), the response fields (spawnedTaskId, seriesEnded), the idempotency scope (one key covers complete, log, and spawn), and the permission model. This is far beyond what a typical description provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient: one long sentence packs the core effect, conditional branches, anchoring rule, response fields, idempotency scope, and permissions. It is front-loaded with the primary action. The sentence is long and could arguably be split for readability, but every clause earns its place and there is no filler. The density is justified by the tool's complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description covers the essential behavioral contract: what happens on completion, what happens for contact-linked tasks, what happens for recurring tasks, how idempotency works, and who is allowed to call it. The only minor gap is that it doesn't describe the return shape beyond spawnedTaskId and seriesEnded, but for a mutation tool with no output schema, the description is remarkably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the schema's silence. It explains the semantics of logActivity (triggers the OP5 timeline entry on contact-linked tasks) and idempotencyKey (covers complete, log, and spawn together). It does not explicitly explain workspaceId or taskId, but those are self-evident from the tool's purpose and the description's context. The description adds meaning for the two non-obvious parameters, which is strong compensation given the 0% coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb and resource ('Erledige eine Aufgabe' / complete a task) and immediately states the core effect: sets done and stamps completedAt. It then enumerates conditional behaviors (contact-linked timeline entry, recurring task spawning) that distinguish it from sibling tools like tasks_cancel, tasks_snooze, and tasks_update. The scope is precise enough that an agent can tell it apart from the other tasks_* siblings without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when the tool applies: completing a task, with conditional branches for contact-linked tasks and recurring tasks. It also states the permission requirement (tasks.write or assignee). However, it does not explicitly say when NOT to use it versus tasks_cancel or tasks_snooze, nor does it name alternatives for non-completion operations. The conditions are implied rather than stated as exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_createA
Lege eine Aufgabe an: a cross-entity to-do with an assignee, an optional due date (dueAt, may be in the past: it lands in the Überfällig bucket), an optional reminder (reminderAt, refused when already past or after the due date), an optional link to any registered record (entityKind/entityId, validated against the OP3 entity registry), and an optional RFC-5545-subset recurrence rule (FREQ=DAILY|WEEKLY|MONTHLY|YEARLY, INTERVAL, BYDAY for weekly, UNTIL or COUNT).
| Name | Required | Description | Default |
|---|---|---|---|
| dueAt | No | ||
| notes | No | ||
| title | Yes | ||
| entityId | No | ||
| entityKind | No | ||
| reminderAt | No | ||
| workspaceId | Yes | ||
| assigneeUserId | Yes | ||
| idempotencyKey | No | ||
| recurrenceRule | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does well. It discloses that dueAt may be in the past and lands in the 'Überfällig' bucket, reminderAt is refused when past or after the due date, entity references are validated against the OP3 registry, and recurrence rules are restricted to an RFC-5545 subset. These are meaningful behavioral details beyond what the schema alone reveals.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence that mixes German and English. Every clause adds value, but the structure could be improved by breaking it into shorter sentences or a list. It starts with the purpose but then becomes a long enumeration of constraints, making it less scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema and no annotations, so the description should clarify return values and overall behavior. It explains key parameter constraints well, but it omits what the tool returns (e.g., created task ID), how idempotencyKey behaves, and any permission or prerequisite context. This leaves an agent partially uninformed for a tool with 10 parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning for dueAt, reminderAt, entityKind/entityId, and recurrenceRule, which is substantial. However, it does not explain notes or idempotencyKey, and workspaceId is only implicit, leaving a few parameters without explicit semantic guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with 'Lege eine Aufgabe an' (create a task), a specific verb and resource, and further defines it as a 'cross-entity to-do'. This clearly states the tool's purpose and distinguishes it from sibling task operations like tasks_list, tasks_update, or tasks_complete, though it does not explicitly name those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the create verb: use this tool when you want to create a new task. However, it provides no explicit guidance on when not to use it or which sibling tools (e.g., tasks_update, tasks_complete, tasks_snooze) should be used for other task operations. No prerequisites or context are given beyond the required fields.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_listARead-only
The task queue (P5): every task with its bucket derived at query time from dueAt against today (overdue, today, upcoming, done), filterable by bucket, assignee, status, or the linked record (entityKind/entityId, which serves the per-entity drawer list). savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| bucket | No | ||
| status | No | ||
| entityId | No | ||
| entityKind | No | ||
| savedViewId | No | ||
| workspaceId | Yes | ||
| assigneeUserId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint=true. The description adds substantial behavioral context: buckets are computed at query time from dueAt, saved view filters are merged underneath explicitly named filters, and entityKind/entityId serves the per-entity drawer list. This provides valuable information beyond the annotation without contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences with every sentence adding distinct value: bucket derivation, filter dimensions, and saved view merging. It is front-loaded with the core purpose and contains no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description thoroughly covers filtering dimensions and saved view behavior, but there is no output schema and no mention of return format, pagination, or sorting. For a list tool, this is a notable gap, although the filter semantics are well described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the bucket parameter (overdue, today, upcoming, done), the savedViewId merge behavior, and the entityKind/entityId use case. However, it does not describe status/assigneeUserId values or workspaceId, leaving some parameter semantics to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: lists the task queue with buckets derived at query time. It clearly distinguishes itself from sibling task tools (tasks_create, tasks_update, etc.) by describing the bucket categories and filter options, making the tool's role unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context on when to use the tool: for the task queue with computed buckets, including the per-entity drawer list scenario. It explains how savedViewId interacts by merging filters, but it does not explicitly name alternatives or exclusions (e.g., tasks_reminders_due), so a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_reminders_dueARead-only
The single reminder-trigger surface (P5): every task whose reminder is due at asOf (default now), status open or doing, and not snoozed past asOf. C01 deal follow-ups, A16 receivables chasing and G06 notification delivery all poll THIS list rather than running reminder loops of their own, and the task.due automation trigger fires from the same predicate.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide only readOnlyHint, so the description takes on the burden of explaining behavior. It adds the selection semantics: asOf defaults to now, only open or doing tasks are included, and snoozed-past-asOf tasks are excluded. This gives the agent a clear model of what will be returned without contradicting the read-only annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core predicate is front-loaded and every sentence serves either purpose or usage. The internal codes (P5, C01, A16, G06) add jargon that may be opaque to an agent, but they do not waste space and the description remains dense and purposeful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only filtered list with two parameters, the description defines the full selection predicate and the default time, while annotations cover the safety profile. It does not describe the return envelope, pagination, or workspaceId semantics, but these are not critical for a correct invocation of this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. It does clarify asOf by giving its meaning and default ('due at asOf (default now)'), but workspaceId semantics are left implicit and additionalProperties is not explained. With only two parameters, this partial compensation is adequate but not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific predicate: every task with a reminder due at asOf, status open or doing, and not snoozed past asOf. It also positions itself as the 'single reminder-trigger surface', clearly distinguishing it from generic task-list tools or custom reminder loops.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names real consumers (C01 deal follow-ups, A16 receivables chasing, G06 notification delivery) and directs them to poll THIS list rather than run their own loops. It also notes the task.due automation trigger uses the same predicate. It stops short of naming sibling alternatives for comparison, so it is not a full when-not-vs-alternative statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_snoozeA
Stelle eine Erinnerung zurück: hides the task from tasks_reminders_due until the given instant. The due date is untouched (snooze hides the reminder, it does not move the deadline), and an instant already past is refused with snooze_in_past.
| Name | Required | Description | Default |
|---|---|---|---|
| until | Yes | ||
| taskId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the main behavior (hides from reminders), what is unaffected (due date), and an error case (snooze_in_past). It doesn't mention permissions or idempotency but covers the primary side effects adequately.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, uses a clear opening verb, and packs essential information into two sentences. It is well-structured and front-loaded, though the mixed German/English phrasing could be cleaner.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 4 parameters and 0% schema coverage, the description is incomplete. It explains the core behavior and error handling but omits parameter details, response format, and any side effects. An agent may struggle to correctly construct the 'until' parameter without further guidance.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining parameters. It does not describe workspaceId, taskId, until, or idempotencyKey at all. Only the concept of 'until' is implied, but no explicit parameter documentation is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (snooze), the resource (task reminder), and the effect (hide from tasks_reminders_due until a given instant). It clearly distinguishes itself from tasks_reminders_due and other task tools by describing its unique behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains when to use the tool (to postpone a reminder) and clarifies what it does not do (does not move the due date) and an error condition (refuses past instants). It doesn't explicitly name alternatives but provides sufficient context for appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasks_updateA
Bearbeite eine Aufgabe from a patch: retitle, reassign, reschedule (patching dueAt clears any snooze: a new deadline supersedes an old Zurückstellen), change the reminder or the recurrence rule, move between open and doing, or reopen a done task (status open; cancelled stays terminal).
| Name | Required | Description | Default |
|---|---|---|---|
| patch | Yes | ||
| taskId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral disclosure burden. It does disclose non-obvious mutation semantics: dueAt supersedes an existing snooze, cancelled tasks cannot be reopened, and done tasks can be reopened to open. It stops short of describing permissions, idempotency behavior, or response shape, but the key task-lifecycle quirks are surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The content is dense and mostly relevant, but it is one long run-on sentence mixing German with English ('from a patch', 'Zurückstellen'), and the parenthetical explanation interrupts readability. Every clause earns its place, yet structure could be tighter and more scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and a nested patch object, the description covers the main operations and the most important status edge cases. It does not document value formats for dueAt, reminderAt, or recurrenceRule, the semantics of patch merge/replace behavior, idempotencyKey usage, or a complete status value list, so an agent still has gaps to resolve before calling correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the meaning of patch fields beyond the bare schema by mapping retitle, reassign, reschedule, reminder, recurrence, and status transitions to the patch object, and adds semantic context for dueAt and status. It omits taskId/workspaceId/idempotencyKey semantics, but those are structurally obvious from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a concrete verb and resource: 'Bearbeite eine Aufgabe' via a patch, and inventories the exact update operations supported (retitle, reassign, reschedule, reminder/recurrence, status moves, reopen). It also differentiates from sibling task tools by noting the reopen rule and terminal cancelled state, so an agent can tell this apart from tasks_complete, tasks_snooze, and tasks_cancel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context for updates and edge-case rules: patching dueAt clears an existing snooze, reopening a done task moves status to open, and cancelled remains terminal. It does not explicitly name sibling tools or state when to prefer tasks_snooze/completed instead, but the behavioral guidance strongly implies the division of labor.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tax_provision_previewARead-only
Propose the Steuerrückstellung for the fiscal year to periodEnd the way Kanton Zürich ZStB 27/1 computes it: profit before tax times s/(1+s) at rateBp (default 2000 = 20 %, the rate is cantonal and yours to set), minus the instalments already charged to 8900, never below zero. Returns the figures and a ready provision_create draft (reason steuern, Dr 8900 / Cr 2330). On an Einzelfirma answers applicable:false. A read: nothing is posted.
| Name | Required | Description | Default |
|---|---|---|---|
| rateBp | No | ||
| periodEnd | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description explicitly states 'A read: nothing is posted', explains the formula and default rate, mentions the edge case for Einzelfirma (applicable:false), and clarifies that it returns a draft. This adds substantial behavioral context beyond the annotation and does not contradict it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-organized in a single paragraph, front-loaded with the purpose, then the formula, output, and edge case. Every sentence contributes information without redundancy. The structure aids quick comprehension.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the computation, parameters, output (figures and a ready draft), edge case, and read-only nature. Even without an output schema, the description tells the agent what to expect. No critical information is missing for a preview tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description fully compensates by explaining rateBp (default 2000, cantonal), periodEnd as the fiscal year end, and the implicit workspace context. It ties the parameters to the computation formula, giving them clear meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the specific verb 'Propose' and resource 'Steuerrückstellung', explains the exact computation method (Kanton Zürich ZStB 27/1), and distinguishes itself from sibling tools by noting it returns a draft and is read-only. This makes it clear what it does and how it differs from provision_create or provision_post.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context that it is for fiscal year end and that it is a preview ('nothing is posted'), which implies it is used before provision_create. However, it does not explicitly name alternatives or state when not to use it. The workflow is implied but not fully spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
time_approveA
Gib eingereichte Zeiteinträge frei (Freigeben): the named submitted entries move to approved, stamped with the approving session actor. All-or-nothing: an unknown id or a non-submitted entry refuses the whole call. Gated on the A24 time.approve capability.
| Name | Required | Description | Default |
|---|---|---|---|
| entryIds | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses key behaviors: all-or-nothing failure, actor stamping, and gating on the A24 capability. It also implies that non-submitted entries cause rejection. It does not mention what happens to already-approved entries or whether the operation is reversible, but the disclosed behaviors are valuable and non-contradictory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the primary action and outcome. Every phrase adds value (submitted entries, all-or-nothing, capability gate), with no fluff. It is efficiently structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has three parameters, no output schema, and no annotations. The description covers the core action and behavior well but fails to explain parameter semantics (workspaceId, idempotencyKey) and does not provide explicit usage guidance or alternatives. It is reasonably complete for the action but lacking in parameter and contextual detail.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It hints at entryIds through 'the named submitted entries' but does not explain workspaceId or idempotencyKey. No parameter descriptions are given in the schema either, leaving workspaceId and idempotencyKey entirely unexplained. This is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (approve) and the resource (submitted time entries), with the outcome of moving to approved and stamping the actor. It is specific enough to distinguish from time_submit and time_lock, but it does not explicitly name or differentiate from the sibling approve_entry, so it is not a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage is implied: it is for approving submitted time entries. However, it does not explicitly state when to use this tool versus alternatives like time_submit or approve_entry, nor does it provide exclusions or conditions beyond the all-or-nothing behavior. This is adequate but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
time_deleteA
Lösche einen Zeiteintrag: allowed while open or submitted only; approved, locked or billed time is an ArG working-time record and refuses with entry_locked.
| Name | Required | Description | Default |
|---|---|---|---|
| entryId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses the refusal error for protected states, which is a key behavioral trait. However, it does not mention other aspects like permanence of deletion, idempotency behavior, or permission requirements, so it is not fully transparent but provides significant context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that front-loads the purpose and immediately states the critical constraint. Every clause adds value, with no redundant information, and it is appropriately concise for a delete operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a delete tool with no output schema and no annotations, the description covers the core purpose and the key usage restriction, which is the main differentiator. However, it omits details about parameters (e.g., how to obtain entryId, purpose of idempotencyKey) and any post-deletion effects, leaving some gaps for an agent that needs complete calling context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 3 parameters with 0% description coverage, and the description contains no parameter explanations. The parameter names (entryId, workspaceId, idempotencyKey) are self-explanatory to some degree, but the description fails to compensate for the absent schema descriptions, offering no additional meaning or usage nuances for these fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with the verb 'Lösche' (delete) and the resource 'Zeiteintrag' (time entry), making the purpose explicit. It clearly distinguishes from siblings like time_update, time_submit, and time_approve, and the constraint about allowed states adds further specificity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states explicitly when deletion is allowed ('while open or submitted only') and when it is not (approved, locked, or billed time), even describing the refusal behavior ('refuses with entry_locked'). This gives clear when-to-use and when-not-to-use guidance, covering the main alternative states without naming sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
time_listARead-only
The timesheet read model (P5): entries filtered by project, user, status, billable, unbilled (everything not yet billed) or a started_at range, plus { totalMinutes, billableMinor } with the money figure derived round-once from each entry snapshotted rate, never stored. This slice (status approved, billable, unbilled) is B02 invoicing input. savedViewId applies a saved view (G00): its stored filters are merged underneath any filter named explicitly here.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | ||
| from | No | ||
| status | No | ||
| userId | No | ||
| billable | No | ||
| unbilled | No | ||
| projectId | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint=true annotation, the description discloses that billableMinor is derived round-once from each entry's snapshotted rate and is never stored, plus the savedViewId merge-underneath precedence semantics. These are non-obvious behavioral traits an agent could not infer from the annotation alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, each earning its place: the filter set and output totals, the business context, and the savedViewId merge behavior. There is no filler or repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter read tool, the description covers the filter semantics, the computed output fields, derivation rules, and the saved-view interaction. With no output schema, describing { totalMinutes, billableMinor } is valuable; however, it omits entry shape beyond totals, pagination behavior, and how multiple filters combine.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description carries the parameter documentation burden: it maps the filter dimensions (project, user, status, billable, unbilled, started_at range, savedViewId) to the schema properties and defines the ambiguous terms 'unbilled' and the saved-view merge behavior. Gaps remain: the required workspaceId is unexplained and the from/to mapping to 'started_at range' is only implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise read-model purpose: timesheet entries filtered by project, user, status, billable, unbilled, or a started_at range, with computed totals. It is clearly distinguished from the mutating time_* siblings (time_start, time_log, time_submit) by being framed as a read model and invoicing input slice.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete usage context: 'This slice (status approved, billable, unbilled) is B02 invoicing input.' This helps an agent know when the tool is relevant, but it does not explicitly name alternatives or state when to prefer a sibling read tool such as billing_unbilled_preview or billing_wip_report.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
time_lockA
Sperre eine Periode (YYYY-MM, optional one project): every approved entry of the period moves to locked, the immutable ArG working-time record billing (B02) reads from. Nothing approved refuses with nothing_to_lock. Gated on the A24 time.approve capability.
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| projectId | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses that every approved entry moves to locked, that the billing record becomes immutable, that the operation fails with 'nothing_to_lock' if nothing is approved, and that it is gated on the A24 time.approve capability. This is rich, non-obvious behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with the action front-loaded; each sentence adds operational value. The phrasing 'Nothing approved refuses with nothing_to_lock' is slightly cryptic, but overall the description is compact and free of filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core invocation details: period format, optional project scoping, side effect, error condition, and permission prerequisite. It does not describe return values (there is no output schema) or mention reversibility via unlock_period, but for a side-effecting state-transition tool the essential context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It usefully defines period as YYYY-MM and marks project as optional, but it leaves workspaceId and idempotencyKey completely unexplained, so two of the four parameters are still dependent on naming conventions rather than documented semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a concrete action ('Sperre eine Periode') with an explicit period format and optional project scope, and the second explains the business consequence: approved entries become locked and feed the immutable billing record B02. It is clear and substantive, but it does not explicitly contrast itself with siblings like lock_period or unlock_period, so sibling differentiation is only implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool by tying it to billing-readiness and states the capability gate ('A24 time.approve'), which gives an entry condition. However, it does not explain when to prefer this over related tools such as lock_period or unlock_period, nor does it state exclusions, so usage guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
time_logA
Erfasse Zeit manuell (nachtragen): a finished entry with startedAt and minutes (1 to 1440), billable by default, priced by the rate card valid on the entry day (OP1 snapshot; no_rate_defined with no card, invalid_minutes outside the bound).
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| userId | Yes | ||
| minutes | Yes | ||
| phaseId | No | ||
| billable | No | ||
| projectId | Yes | ||
| startedAt | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and discloses important behaviors: billable by default, pricing via rate card valid on entry day (with OP1 snapshot), and edge cases (no_rate_defined, invalid_minutes). However, it omits details like idempotency or creation side effects, so it is not fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence that front-loads the purpose and packs in key constraints and pricing rules. It is concise without being overly long, though the mix of German and English could be seen as slightly awkward but still efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 9 parameters (5 required) and no output schema, the description is incomplete. It omits explanations for most parameters, the purpose of idempotencyKey, and any information about the return value or error handling. An agent would need more guidance to call this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains startedAt, minutes (with the 1-1440 constraint), and billable default, but fails to cover other required parameters like workspaceId, userId, projectId, and optional ones like phaseId, notes, idempotencyKey. The compensation is partial at best.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states this tool is for manually recording time entries retroactively ("Erfasse Zeit manuell (nachtragen)"), specifying key fields (startedAt, minutes) and the default billable behavior. It distinguishes itself from live tracking tools like time_start/time_stop by emphasizing the manual retroactive nature.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'manuell (nachtragen)' implies usage for retroactive entry as opposed to live timers, but it does not explicitly name alternatives or provide exclusion criteria. An agent could infer when to use it, but explicit routing to siblings like time_update or time_start is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
time_resolve_rateARead-only
Resolve which hourly rate WOULD price time for a user/project/client at a date (OP1, the single resolver): answers the winning card (rateMinor, currency, sourceScope, rateCardId) by precedence client, project, employee, default, or no_rate_defined when no card is valid at that day.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | ||
| userId | No | ||
| contactId | No | ||
| projectId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the operation read-only, and the description adds genuine behavioral detail: it returns a winning rate card or no_rate_defined, and follows a defined precedence chain. No contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense, front-loaded sentence that efficiently packs purpose, outputs, precedence, and fallback behavior. The heavy use of parenthetical jargon (OP1, rateMinor, sourceScope) slightly reduces readability but does not waste words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 5 parameters, no output schema, and no parameter documentation, the description covers the main behavior, return shape, and fallback well. But it omits the required workspaceId context and the format of at, and leaves rateMinor/currency/sourceScope undefined, so the description is not fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description partially compensates by linking client/project/employee/date to the relevant parameters and precedence. It still leaves workspaceId unexplained and gives no format or constraint detail for at, so an agent has to infer some parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Resolve which hourly rate would price time'), identifies the resource scope, lists the exact output fields, and gives the precedence rule. It clearly differentiates itself as 'the single resolver' from ambiguous price-related siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the resolver's behavior and precedence explicit, so an agent can infer when this tool is relevant. However, it never names alternatives like price_resolve or rate_card_list, nor gives an explicit 'use this when...' or 'use that when...' rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
time_startB
Starte einen Timer auf einem Projekt (B00): creates an open time_entry with started_at now and the rate snapshotted through resolveRate (OP1: client, then project, then employee, then default card). One running timer per user (timer_already_running names the running entry); with no valid rate card it refuses with no_rate_defined rather than minting a 0-rate entry.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| userId | Yes | ||
| phaseId | No | ||
| billable | No | ||
| projectId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly states that it creates a time entry with a snapshot of the rate via resolveRate, enforces a single running timer per user, and refuses with no_rate_defined instead of creating a zero-rate entry. This provides meaningful behavioral insight beyond the basic action.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence but is appropriately structured: it opens with the core action ('Starte einen Timer auf einem Projekt') and then provides crucial constraints and failure modes. It is concise but packed with necessary details, though it could benefit from better front-loading of the main purpose over the rate resolution details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (rate resolution, concurrency constraints) and the lack of annotations or output schema, the description is incomplete. It explains the rate resolution and the one-timer rule, but omits the return value, does not clarify the required parameters, and ignores idempotency and other optional fields. An agent would need to infer parameter usage from the schema alone, which has no descriptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, meaning the description must compensate for undocumented parameters. It does not explain any of the seven parameters, including required ones like workspaceId, userId, and projectId, nor does it mention optional fields such as notes, billable, phaseId, or idempotencyKey. The description fails to add any parameter-level meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action: starting a timer on a project, creating an open time_entry. It distinguishes itself from related tools like time_stop and time_log by focusing on the start action and mentioning the rate resolution logic, which clarifies its unique role.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives context about one running timer per user and the refusal condition when no rate card exists, which helps an agent understand when it might fail. However, it does not explicitly compare with alternatives like time_stop or time_log, nor does it state prerequisites such as the need for an existing project.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
time_stopA
Stoppe den laufenden Timer: stamps ended_at and computes the full elapsed minutes, midnight crossings included (the entry keeps its started_at day; day-splitting is a report concern). The entry stays status open and editable.
| Name | Required | Description | Default |
|---|---|---|---|
| entryId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It goes beyond the name by explaining that it stamps ended_at, computes full elapsed minutes including midnight crossings, preserves the original started_at for day attribution, and leaves the entry 'status open and editable'. This is substantial, though it omits details like idempotency or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight sentences, with the core action front-loaded and supplementary behavioral details appended without redundancy. Every phrase earns its place: the action, the computation, the day-keeping nuance, and the status outcome. It is concise yet highly informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and no annotations, the description provides essential operational context (elapsed minutes, day-keeping, status). However, it does not clarify parameter semantics, failure scenarios, return values, or the role of idempotencyKey. These gaps leave the agent with incomplete information for calling the tool correctly in all situations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 3 parameters (entryId, workspaceId, idempotencyKey) with 0% description coveragechend. The description does not explain what these parameters mean or how they map to the tool's behavior; it only refers to 'the entry' indirectly. Since schema coverage is zero, the description must compensate but fails to do so, leaving the agent to infer parameter roles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Stoppe den laufenden Timer' (stop the running timer) and specifies it 'stamps ended_at and computes the full elapsed minutes'. This is a specific verb+resource combination that distinguishes it from sibling tools like time_start and time_log, though it does not explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the usage condition: it stops a 'laufenden' (running) timer, so it should be used when a timer is in progress. However, it does not explicitly mention when not to use it or draw contrasts with alternatives, so the guidance remains implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
time_submitA
Reiche die Zeit einer Periode ein (YYYY-MM, optional one project): every finished open entry of the period moves open to submitted; a period with nothing open refuses with nothing_to_submit. A still-running timer is not swept up.
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| projectId | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does a strong job: it discloses the state transition, the error condition when nothing is open ('refuses with nothing_to_submit'), and the important exclusion of still-running timers. It does not cover every side effect, such as permission requirements or what happens to unfinished open entries explicitly, but it is far beyond a bare mutation description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the primary action, then packs in the key behavioral conditions without fluff. The mixing of German and English is slightly awkward and could reduce clarity for non-German readers, so it is not a perfect 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description covers the essential call-time context: period format, optional project, state transition, the nothing_to_submit error, and the running-timer exclusion. The main omission is semantic guidance for the idempotencyKey parameter, but for a moderately complex mutation tool the description is largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It valuably documents the period format (YYYY-MM) and the optional project scoping, which maps to the period and projectId parameters. However, it does not clarify workspaceId or idempotencyKey semantics, leaving two parameters effectively unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action on a resource: submitting time for a period, with a precise state transition ('moves open to submitted'). It clearly distinguishes itself from sibling time tools like time_start, time_stop, time_log, time_update, and time_approve by describing a period-level submission workflow rather than individual entry editing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context: it applies to a period (YYYY-MM), can be scoped to one optional project, and only affects finished open entries. However, it does not explicitly state when to prefer this tool over closely related alternatives such as time_approve or time_lock, and it does not provide any when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
time_updateB
Bearbeite einen Zeiteintrag from a patch (minutes, billable, notes, startedAt, projectId, phaseId): allowed while open or submitted; from approval onward the record is frozen and refuses with entry_locked. Re-pointing to another project re-resolves the rate snapshot at the entry own capture day; a later rate-card edit never reprices captured time.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | Yes | ||
| entryId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Given no annotations, the description carries full responsibility and it does disclose key behaviors: the entry_locked refusal after approval and the rate snapshot behavior when re-pointing projects. This adds substantial context beyond the basic 'edit' action, though it does not cover idempotency or other edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and information-dense. It front-loads the primary action and patch fields, then adds crucial behavioral constraints in a clear sequence. It earns every phrase, though the 'from a patch' phrasing is slightly awkward.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the state machine and rate behavior, which are critical contextual details. However, it omits parameter semantics and does not address expected response or error handling beyond entry_locked.src. For a mutation tool with no annotations, a complete description should also clarify parameter formats and idempotency behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. It lists patch subfields (minutes, billable, notes, etc.) but only repeats their names without explaining meaning, formats, or constraints. It also ignores the top-level parameters (workspaceId, entryId, idempotencyKey), leaving significant gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it edits a time entry ('Bearbeite einen Zeiteintrag') and enumerates the patchable fields, which is specific and understandable. However, it does not explicitly distinguish itself from sibling tools like time_delete or time_log, so it lacks direct differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides concrete conditions for when the tool is allowed (open or submitted states) and when it fails (from approval onward), which informs usage. It does not name alternative tools or state explicitly when to use a sibling instead, so the guidance is implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
transition_documentC
Advance a document (issue, send, accept, decline, confirm, cancel); issuing an invoice posts via the delegate.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | ||
| documentId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It does note that issuing an invoice 'posts via the delegate,' which adds a useful implementation detail, but it does not disclose side effects, reversibility, permissions, or consequences of transitions. This is significant for a mutation-oriented tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single compact sentence with the core action front-loaded and the transition list immediately following. The closing clause about the delegate is cryptic but still adds unique behavioral information without unnecessary padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a 4-parameter, mutation-facing tool with no annotations, no output schema, and 0% schema parameter coverage. The description does not explain valid values for 'to,' idempotency semantics, or any preconditions. It is far from sufficient for an agent to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description provides no meaning for workspaceId, documentId, to, or idempotencyKey. It does not explain that 'to' likely indicates the target state or how the listed actions map to it. The description fails to compensate for the schema's lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('advance') with a clear resource ('document') and enumerates the supported transition actions (issue, send, accept, decline, confirm, cancel). This conveys the general purpose well, though it does not explicitly differentiate itself from sibling tools like send_invoice or quotes_accept.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical list implies which lifecycle actions are supported, but there is no guidance on when to use this generic transition tool versus the more specific sibling tools (e.g., send_invoice, quotes_accept, quotes_decline). No alternative tools or exclusions are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
trial_balanceARead-only
Saldenbilanz (trial balance): every account's opening balance, its debit and credit movement in the period, and its closing balance, in the workspace base currency and integer Rappen. Figures are debit-positive, exactly as the ledger holds them, and drafts are excluded. Pass compareTo={periodStart,periodEnd} for the prior-period column and its delta. reconciles reports three checks by name (debit equals credit, the closing column ties to an independently queried ledger balance, and every account the ledger moved was rendered); it is evidence about coverage and the period boundaries, not proof that an account sits in the right section. groupBy accepts 'kmu' today and refuses anything else, because account-level custom fields are a G00 capability that does not exist yet.
| Name | Required | Description | Default |
|---|---|---|---|
| groupBy | No | ||
| compareTo | No | ||
| periodEnd | Yes | ||
| periodStart | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses important behavior: debit-positive convention, draft exclusion, currency/Rappen handling, reconciliation semantics, and the current groupBy limitation. These details give the agent an accurate model of the tool's behavior without overpromising.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence adds operational value: report contents, sign convention, exclusions, comparison parameter, reconciliation caveats, and groupBy constraint. It is front-loaded with the core meaning and then covers edge semantics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
In the absence of an output schema, the description effectively conveys what the response contains and important caveats. It could be slightly more explicit about the output container or pagination, and does not mention the optional 'source' property inside compareTo, but overall it is sufficient for safe invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates well for compareTo and groupBy, explaining their purpose and accepted values. However, workspaceId and the exact date format for periodStart/periodEnd are left implicit, though inferable from context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly specifies the tool's output: every account's opening balance, debit/credit movements, and closing balance. It differentiates this from sibling financial reports by emphasizing account-level detail, base currency, and integer Rappen.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides concrete usage context: compares with a prior period via compareTo, restricts groupBy to 'kmu', and clarifies what the reconciliation checks do and do not prove. It does not explicitly name sibling alternatives or state when not to use this tool, but the parameter-level guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unarchive_accountC
Reactivate a soft-archived account (idempotent).
| Name | Required | Description | Default |
|---|---|---|---|
| accountId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral burden. It only discloses that the operation is idempotent. It does not mention permissions, side effects, error behavior, or what happens if the account is already active.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is very concise (one sentence), but it is under-specified. While there is no fluff, the brevity leaves out essential information, so it is not effective conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema and no annotations, the description should explain what happens on success/failure, return values, and any constraints. It only mentions idempotence, leaving many operational details missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description provides no parameter details. The parameters (workspaceId, accountId) are self-explanatory from their names, but the description adds zero value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (reactivate) and resource (account), and mentions 'soft-archived' to indicate the state. It implies it is the inverse of archive_account but does not explicitly name that sibling, so it slightly misses differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It says 'Reactivate a soft-archived account' which implies usage for archived accounts, but it does not explicitly state when to use it versus alternatives like delete_account or archive_account, nor mention any prerequisites or context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unarchive_bank_accountA
Wiederherstellen: put an archived Bankkonto back in the pickers. The inverse of archive_bank_account, and a Bankkonto that is already active stays active (no rejection). Nothing about the account changes: this only clears the archived flag.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| bankAccountId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: it discloses no rejection for already-active accounts, no other state changes, and that only the archived flag is cleared. This is unusually explicit about side effects and non-destructiveness.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each adding unique information: the restore action, the idempotent already-active behavior, and the precise side-effect scope. There is no filler or repeated schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-required-field mutation with no output schema, the description fully explains purpose, behavior, and side-effect boundary. An agent has enough to invoke it correctly and know what will happen.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description provides no parameter-specific guidance. workspaceId, bankAccountId, and idempotencyKey are only self-descriptive by name, and the description does not compensate for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description defines a specific restore action on a Bankkonto, names the inverse archive_bank_account, and states the sole effect (clears the archived flag). This clearly distinguishes it from account-level archive/unarchive tools and other bank account operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The inverse-of-archive statement and the explicit note that already-active accounts stay active give clear context for when to use the tool, including the idempotent edge case. It does not explicitly name alternatives or exclusions such as unarchive_account for ledger accounts, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unarchive_contactA
Reactivate a soft-archived contact (idempotent).
| Name | Required | Description | Default |
|---|---|---|---|
| contactId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses idempotency, a useful behavioral trait. However, with no annotations provided, it does not cover permissions, side effects, or behavior when the contact is not already archived. The description carries the full burden but only provides minimal behavioral information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence with no filler. The action is front-loaded, and the idempotency note is efficiently placed. It earns its place without unnecessary detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple toggle operation, the description covers the core action and idempotency. However, given the absence of annotations and an output schema, it lacks details on prerequisites (e.g., contact must be archived) and potential side effects. It is adequate but not fully complete for an agent to call it confidently without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not explain the parameters (workspaceId, contactId) at all. The schema only lists them as strings with no descriptions. Since the description must compensate for low schema coverage, its failure to provide any parameter guidance is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Reactivate a soft-archived contact' with the verb 'Reactivate' and resource 'contact'. It also notes idempotency, making it distinct from archive_contact and other contact operations. The purpose is unambiguous and differentiates from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used to undo archiving, which is clear from the context. It does not explicitly mention alternatives or when not to use it, but for a simple unarchive operation the usage context is self-evident. There is no mention of prerequisites, but the action is specific.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unarchive_cost_centerB
Reactivate a soft-archived cost centre (idempotent).
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| costCenterId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must carry the behavioral burden. It usefully discloses that the operation is 'idempotent' and limited to 'soft-archived' cost centres, which are meaningful behavioral traits. However, it does not mention side effects, permission needs, or behavior on already-active cost centres, so transparency is partial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence that communicates the action, target state, and idempotency without any filler words. It is appropriately sized for the simplicity of the operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter lifecycle operation, the description provides the essential purpose and idempotency note, and the schema covers parameter requirements. It is complete enough for basic invocation but misses guidance on related tools (archive/delete) and any side-effect clarification, which would make it fully self-sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate by explaining the parameters. It does not mention workspaceId or costCenterId at all; while the names are self-explanatory, the description adds no semantic value beyond what the schema already shows.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly specifies the verb 'reactivate' and the resource 'soft-archived cost centre', making the operation unambiguous and distinct from archive_cost_center and delete_cost_center. It does not explicitly name sibling tools, so it stops short of full differentiation, but the verb+resource identification is specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case (restoring a previously archived cost centre) but provides no explicit guidance on when to use it vs alternatives like archive_cost_center or delete_cost_center. There are no exclusions, prerequisites, or decision rules to help an agent select this over related lifecycle tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unarchive_itemA
Reactivate a soft-archived item (idempotent).
| Name | Required | Description | Default |
|---|---|---|---|
| itemId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full disclosure burden. It does add the valuable trait of idempotency and clarifies that the item is in a soft-archived state, which partially describes behavior. However, it does not disclose permissions, potential side effects on related data, or what a successful reactivation changes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler; every word earns its place because 'reactivate' states the action, 'soft-archived' states the target, and 'idempotent' adds a key behavioral constraint.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-string-parameter reactivation tool, the description covers the core operation and idempotency, which is mostly sufficient at a basic level. However, with no annotations and no output schema, it omits prerequisites, success criteria, and edge cases, leaving gaps an agent would need to infer.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description says nothing about workspaceId or itemId, so the agent must infer from the parameter names that itemId identifies the item and workspaceId scopes the operation. The names are self-explanatory, but the description adds no semantic value beyond the schema's bare property names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('reactivate') with a specific resource ('soft-archived item'), and the 'soft-archived' qualifier clearly distinguishes it from hard-delete or recreate operations. The sibling archive_item makes the inverse relationship obvious without needing to inspect the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context implies this tool is the inverse of archive_item and should be used on soft-archived items, but the description never explicitly states when to use it, what prerequisites apply, or how to choose it over similar reactivation tools like unarchive_account or unarchive_contact. Usage guidance is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
uninstall_pluginA
Deinstalliere eine Erweiterung (US-G02.3): disable it first (sweep its registrations), record the removal in the append-only audit_log before the row is gone, then delete the manifest row (plugin manifests are not §H-AUDIT financial data, so a real delete is correct). Traceability rides the audit chain, not an append-only manifest table.
| Name | Required | Description | Default |
|---|---|---|---|
| pluginId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure, and it does so thoroughly. It reveals destructive behavior (deleting the manifest row), explains the required ordering (disable first, audit-log before deletion), and gives policy context for why a real delete is correct versus an append-only approach. This is rich, honest behavioral information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but mostly earns its length by explaining the required deletion sequence, audit logging, and the compliance rationale. The reference codes (US-G02.3, §H-AUDIT) add context but could be trimmed without losing clarity. It is structured with the core action first, followed by the procedural detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description does a strong job on process and audit semantics, but it omits the return behavior/output, does not explain the idempotency parameter, and leaves "sweep its registrations" somewhat vague. Given that this is a destructive tool with no annotations and no output schema, these are meaningful gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description provides no explanation of workspaceId, pluginId, or idempotencyKey. The idempotencyKey, in particular, is a nontrivial parameter whose semantic role is entirely undocumented. With no parameter information in the schema and none in the description, an agent cannot map inputs to the tool's behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: "Deinstalliere eine Erweiterung" (uninstall an extension). It further clarifies what uninstalling entails versus merely disabling by describing the sequence: disable first, log to audit_log, then delete the manifest row. This clearly distinguishes it from siblings like disable_plugin and install_plugin.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies that uninstall is the permanent-removal counterpart to disable, and it states that disabling must happen first. However, it never explicitly names alternatives such as disable_plugin or states when to choose one over the other. Usage context is inferable but not directly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unlock_periodA
Unlock a period (a filing/year seal refuses). CONSEQUENCE: Lifts a soft period lock, so postings into a period a human had locked are accepted again; a hard seal refuses.
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the key behavioral consequence: lifting a soft period lock so postings are accepted again, and explicitly states that a hard seal refuses. This is meaningful behavioral context beyond the schema. It doesn't mention idempotency or side effects on audit logs, but the core mutation behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the action and consequence. The 'CONSEQUENCE:' label is a useful structural cue. It earns its place, though the parenthetical '(a filing/year seal refuses)' is slightly awkward and could be clearer.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no annotations and no output schema, the description covers the core behavior and the hard-seal limitation. However, it doesn't explain the idempotencyKey parameter, which is required and unusual, nor does it describe what happens if the period is already unlocked or hard-sealed. Given the tool's simplicity, this is adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain any of the three parameters (workspaceId, period, idempotencyKey). The description mentions 'period' conceptually but doesn't clarify its format or that it's the target of the unlock. With 0% coverage and no param details, the description fails to compensate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Unlock') and resource ('a period'), and clarifies the context (a filing/year seal refuses). It distinguishes the tool from lock_period and close_year/reopen_month siblings by focusing on the soft-lock vs hard-seal distinction. However, it doesn't explicitly name a sibling alternative, so it's clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when a period is soft-locked and postings need to be accepted again. It contrasts with a hard seal that refuses, which tells the agent this tool won't work on hard-sealed periods. It doesn't explicitly name alternatives like lock_period or list_period_locks, but the context is clear enough for an agent to decide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_accountA
Rename an account or change its VAT default / cost-centre flag (number and type are frozen).
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| accountId | Yes | ||
| workspaceId | Yes | ||
| vatCodeDefault | No | ||
| costCenterAllowed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that the operation is a mutation ('Rename' / 'change') and that certain fields are frozen ('number and type are frozen'), which is useful behavioral context. However, it does not disclose whether the change is reversible, whether it affects existing references (e.g., historical transactions), or whether any validation occurs (e.g., duplicate name checks). For a mutation tool with no annotations, this is a moderate gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that front-loads the primary action ('Rename an account') and then efficiently lists the secondary capabilities and the frozen constraints. Every word earns its place; there is no fluff or repetition of schema information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter mutation tool with no annotations and no output schema, the description is reasonably complete: it states the purpose, the mutable fields, and the frozen fields. However, it does not mention return values, error conditions, or side effects (e.g., whether renaming affects historical records or requires specific permissions). Given the tool's complexity (5 params, mutation, no annotations), a bit more context would be needed for an agent to call it with full confidence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does add meaning by mapping the purpose to the parameters: 'name' (rename), 'vatCodeDefault' (VAT default), 'costCenterAllowed' (cost-centre flag). However, it does not explain the required parameters 'workspaceId' and 'accountId' (though these are self-evident as identifiers), nor does it clarify the format or constraints of 'vatCodeDefault' or the exact semantics of 'costCenterAllowed'. The description adds some value but leaves the agent to infer parameter details from names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Rename') and resource ('an account'), and immediately clarifies the scope of what can change: 'or change its VAT default / cost-centre flag'. It also explicitly states what is frozen ('number and type are frozen'), which distinguishes this from other account-related tools like create_account, archive_account, or update_company_profile. This is a clear, specific, and differentiating description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use this tool: when you need to rename an account or change its VAT default / cost-centre flag. It also implicitly excludes other operations by stating what is frozen (number and type). However, it does not explicitly name alternative tools (e.g., create_account for new accounts, archive_account for deactivation) or state when NOT to use it. The context is clear but exclusions are not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_automation_ruleA
Patch a rule: its name, trigger, condition or action. The patched shape is validated exactly as a new rule is, so an edit can never leave a rule the engine would have refused to create. The change takes effect on the next event; an archived rule refuses.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | Yes | ||
| ruleId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden and does well: it discloses validation behavior (patched shape validated exactly as a new rule), timing (takes effect on next event), and rejects archived rules ('an archived rule refuses'). This gives an agent critical behavioral expectations beyond what the schema shows.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all valuable. The core action and patched fields are front-loaded, followed by validation semantics and timing/archive behavior. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool involves a nested patch object and 4 params with no output schema, the description conveys the key context: what can be patched, validation equivalence, effect timing, and archive refusal. It doesn't explain the patch object's exact shape (e.g., which fields are required inside the patch), which would help an agent construct the call, but the critical operational semantics are covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needs to compensate. The description names the patched fields (name, trigger, condition, action) which maps to the 'patch' object parameter, and implies ruleId is the target and workspaceId scopes it. However, it doesn't detail the patch structure or idempotencyKey semantics, leaving significant parameter meaning unexplained. Baseline for 0% coverage would be lower; naming the patch dimensions helps but leaves gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the tool patches an automation rule and specifies the patched dimensions: name, trigger, condition, or action. This clearly distinguishes it from sibling tools like create_automation_rule, enable_automation_rule, and archive_automation_rule.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool: to patch existing rules, noting that edits take effect on the next event. It implicitly distinguishes from creation by emphasizing validation exactly as a new rule. It could be stronger by explicitly naming alternatives like create_automation_rule for new rules, but the contrast with engine-creation validation makes the boundary clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_bank_accountA
Rename a Bankkonto or correct its details. The IBAN, currency and verknüpftes Konto freeze once the account is referenced by a posted opening balance (account_in_use); the name stays editable for good.
| Name | Required | Description | Default |
|---|---|---|---|
| iban | No | ||
| name | No | ||
| currency | No | ||
| workspaceId | Yes | ||
| bankAccountId | Yes | ||
| idempotencyKey | No | ||
| ledgerAccountId | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full behavioral disclosure burden. It does disclose a significant state-dependent behavior: IBAN, currency, and verknüpftes Konto freeze after the account is referenced, while the name remains editable. However, it does not explain error behavior when attempting to update frozen fields, permission requirements, idempotency semantics, or what the response contains. The disclosure is partial but valuable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The first sentence delivers the core purpose immediately, and the second adds the most important operational constraint. Every word contributes either to clarifying the action or to preventing a failed call.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 7 parameters, no output schema, and no annotations, the description is not complete enough for an agent to safely invoke it. It lacks details on the return value, error conditions for frozen fields, the meaning of ledgerAccountId/idempotencyKey, and any conditional requirements. The freeze rule is helpful, but too many critical invocation aspects remain undisclosed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning to several parameters by mapping them to business concepts: IBAN, currency, and verknüpftes Konto (likely ledgerAccountId) are called out as freezable, and name is described as always editable. However, it leaves idempotencyKey, workspaceId, and bankAccountId without explanation, and the mapping of 'verknüpftes Konto' to a specific schema field is indirect.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb phrase 'Rename a Bankkonto or correct its details', which clearly identifies the action and resource. The resource is specific to bank accounts, distinguishing it from generic account tools like update_account, and the 'or correct its details' phrasing covers the full update scope without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used to update bank account fields, and it provides a crucial conditional rule (fields freeze once the account is referenced by a posted opening balance). However, it does not explicitly state when to prefer this tool over siblings such as create_bank_account or update_account, nor does it define exclusions. The guidance is inferred from the purpose rather than explicitly framed as usage direction.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_company_profileB
Update company identity fields (name, legal form, UID, MWST number).
| Name | Required | Description | Default |
|---|---|---|---|
| uid | No | ||
| name | No | ||
| mwstNo | No | ||
| legalForm | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing behavior, but it only says 'Update', which is already implied by the tool name. It does not explain whether it performs a partial or full update, what happens to unspecified fields, whether authorization or specific permissions are needed, or any side effects. This is a significant gap for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler words. Every word adds value, efficiently conveying the tool's purpose and the key fields involved.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters, 0% schema coverage, no annotations, and no output schema, this description is incomplete. It fails to mention the required workspaceId, the update semantics (partial vs. full), what happens if a field is omitted, or any expected return value. An agent would need to guess at critical usage details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, so the description must compensate. It does add some meaning by mapping the fields (e.g., 'MWST number' to mwstNo, 'legal form' to legalForm), which helps an agent understand the purpose of these parameters. However, it omits the workspaceId parameter entirely, which is the only required parameter, and provides no format or validation details for the optional fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Update'), a clear resource ('company identity fields'), and enumerates the specific fields involved (name, legal form, UID, MWST number). This clearly distinguishes it from the related sibling tool get_company_profile and other profile-related tools like set_creditor_profile.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as get_company_profile or set_creditor_profile. It does not mention any prerequisites (e.g., workspaceId requirement) or scenarios where this update is appropriate versus another tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_contactC
Update a contact from a patch.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | Yes | ||
| contactId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full burden. It only states the action and the patch mechanism, but does not disclose whether it performs a full replace or partial merge, what happens if the contact does not exist, whether it triggers side effects like audit logging, or what the response contains.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with the action front-loaded. However, it is under-specified for the tool's complexity, so while it is not verbose, it lacks necessary detail to be considered well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a nested patch object, no annotations, no output schema, and no parameter descriptions. The one-sentence description is grossly insufficient for an agent to correctly construct a patch or understand the update behavior. It fails to cover essential context like field constraints, error conditions, or return values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It fails to explain the meaning of workspaceId and contactId, or what the 'patch' object should contain beyond the schema's property names. No parameter semantics are provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Update a contact from a patch' clearly states a specific action (update) on a specific resource (contact), and the 'from a patch' qualifier indicates a partial/merge semantics. It distinguishes from sibling contact tools like create_contact, archive_contact, get_contact, and list_contacts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives. It does not mention prerequisites (e.g., contact must exist), when to prefer create_contact over update, or any conditions that would make this tool inappropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_documentA
Patch a draft document (draft-only; an issued document is immutable and corrected by reversal).
| Name | Required | Description | Default |
|---|---|---|---|
| patch | Yes | ||
| documentId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of disclosing behavior, and it does disclose the key mutation semantics: it patches drafts, and issued documents are immutable and corrected by reversal. It does not mention partial-patch behavior, permissions, or error cases, but it covers the critical lifecycle constraint an agent needs to avoid misuse.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence with high information density. The primary action is front-loaded ('Patch a draft document'), and the caveat about issued documents is added without filler. Every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has four parameters including a deeply nested patch object and no output schema, so an agent needs more than a one-line lifecycle statement. The description leaves patch mechanics, idempotencyKey purpose, required fields within nested lines, and return behavior unexplained. It is too thin for a mutation tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it does not mention any parameter. The important patch semantics, the purposes of idempotencyKey, and the meaning of the nested line fields are left entirely to schema inference. The word 'Patch' only weakly implies partial update.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'Patch' and names the resource 'a draft document', immediately distinguishing it from create/convert/transition document tools. It adds explicit lifecycle context ('draft-only') and the immutability rule for issued documents, so an agent can tell which operation this is without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when to use it: on draft documents only. It also gives an exclusion ('an issued document is immutable') and points to reversal as the correct correction path, which is useful routing guidance. It does not name sibling tools explicitly, but the draft-versus-issued distinction is enough context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_document_templateA
Update a template (name, footer, language, optional columns, logo). NEVER changes an already-issued document: issue froze the template content onto the document, so a mailed copy and a later reprint always match; only documents issued after the edit pick it up.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | Yes | ||
| templateId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears the burden of disclosing behavior, and it does: it reveals the non-obvious freeze-on-issue behavior and that prior documents never change. It doesn't mention patch merge semantics or permissions, but the most dangerous misconception—that updating a template retroactively changes sent documents—is explicitly prevented.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler: the first gives the action and field scope, the second delivers the critical caveat. The warning is front-loaded and every phrase contributes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an update tool with no annotations, no output schema, and a nested patch object, the description covers the essential side-effect and the editable scope. It falls just short of specifying patch application behavior and retry/idempotency handling, but the core invocation is well supported.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and patch is an opaque object, so the description adds value by enumerating the editable fields (name, footer, language, optional columns, logo). It does not explain how the patch is applied (merge vs replace) nor the roles of workspaceId, templateId, and idempotencyKey beyond their names, so the compensation is only partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Update), resource (template), and editable scope (name, footer, language, optional columns, logo). It also draws a clear behavioral line against update_document by explaining that already-issued documents are untouched, so the tool's role is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for when the tool applies: edits affect only documents issued after the update, not already-issued ones. It does not explicitly name sibling alternatives like create_document_template or archive_document_template, but the update verb and freeze-on-issue semantics make the intended use obvious.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_itemD
Update an item from a patch.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | Yes | ||
| itemId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It only says 'update' which implies mutation, but it does not state whether this is a partial update, what happens to unspecified fields, whether the operation is reversible, or if any permissions are required. The phrase 'from a patch' hints at partial update but is not explicit, leaving significant behavioral traits undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise at one short sentence, but this is under-specification rather than effective conciseness. It is front-loaded with the verb and object, but the lack of any additional context makes it inadequate for a tool with a nested patch object and three required parameters. It is too terse to be useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (three required parameters including a nested object, no annotations, no output schema), the description is severely incomplete. An agent cannot determine what a patch is, what fields are updatable, any constraints, or the effect of the update. It fails to provide essential context for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, and the description adds no information about the parameters. It does not explain what 'patch' means, what fields are updatable, or any constraints (e.g., optionality, allowed values). The agent must rely solely on the schema, which has no descriptions for the properties, so the parameter semantics are essentially undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'update' and the resource 'item', so it is not a tautology. However, the phrase 'from a patch' is ambiguous and does not clarify the scope or distinguish it from sibling operations like create_item or delete_item beyond the name itself. It lacks any mention of what aspects can be updated or that it is a partial update, leaving the purpose only partially clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance whatsoever on when to use this tool versus alternatives. The description does not mention prerequisites, conditions, or situations where update_item is preferred over other item-related tools (e.g., create_item, archive_item). The agent is left to infer usage entirely from the name and schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_recurring_scheduleA
Patch a schedule: template, cadence, bounds or autoIssue, as absolute values. Changes affect only FUTURE generations; already produced invoices are A10/A11 documents and stay untouched. A cadence change (interval, customDays, anchorDate) restarts the series at the next occurrence on or after today. An ended schedule refuses with schedule_ended; an end date before the anchor refuses with end_before_anchor.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | Yes | ||
| scheduleId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses that changes affect only future generations, cadence changes restart the series, and it refuses under specific error conditions. It could go further by mentioning whether the operation is reversible or any side effects, but it covers the most critical behaviors.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded: it states the action, scope, and key behavioral constraints in three sentences. Every sentence adds value—no fluff. The key constraints (future-only, cadence restart, error conditions) are clearly prioritized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (nested patch object) and lack of schema descriptions, the description is remarkably complete. It covers what the tool does, how it affects schedules, and error conditions. It mentions the patch fields and the behavior for cadence changes. An agent has enough to call this correctly, though the exact patch structure isn't fully defined (but that's beyond the description's scope if the schema is available).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only parameter names and types, with no descriptions (0% coverage). The description explains what the patch object can contain (template, cadence, bounds, autoIssue) and how cadence changes behave. It adds significant meaning beyond the schema, though it doesn't detail the exact structure of the patch object beyond listing field names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool patches a schedule with specific fields (template, cadence, bounds, autoIssue) as absolute values. It distinguishes this from other schedule operations like pause, resume, or end by emphasizing it modifies future generations and has specific error handling. The verb 'patch' combined with 'schedule' makes the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly explains when to use this tool: for patching a schedule's properties. It also gives critical usage context: changes only affect future generations, cadence changes restart the series, and the tool refuses under certain conditions. This is clear guidance on when and how to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_saved_viewC
Patch a saved view. A personal view may only be changed by the session that owns it; a workspace-shared one, and publishing a personal one, require manage_saved_views.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | Yes | ||
| viewId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It does disclose permission and ownership constraints: personal views can only be changed by the owning session, while workspace-shared views and publishing require manage_saved_views. However, it does not mention idempotency behavior, error conditions, or whether changes are reversible. The permission context is useful but the description is sparse beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, consisting of two sentences with no filler. The primary action is front-loaded, and the permission note is informative. It is efficient and does not waste words, though it is brief in content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of annotations, output schema, and parameter descriptions, the description is insufficient for an agent to call the tool correctly. It does not explain the structure of the patch object, the purpose of idempotencyKey, or any expected response. The permission info is helpful, but the overall context is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameters. It does not mention any of the four parameters (patch, viewId, workspaceId, idempotencyKey). It only implies 'patch' is a change object but does not specify its structure or fields. The description adds virtually no meaning beyond the schema's bare types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action 'Patch a saved view' with a specific verb and resource. It distinguishes from sibling tools like create_saved_view and delete_saved_view by using 'patch' to imply modification. However, it does not explicitly differentiate from other saved-view tools, so it is clear but not fully distinctive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as create_saved_view or list_saved_views. It only mentions permission requirements for different view types, which is authorization context rather than usage selection guidance. No exclusions or alternatives are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_annual_reconciliationARead-only
The two annual MWST reconciliations of a calendar year as figures (Art. 128 Abs. 2 and 3 MWSTV, the ESTV's Umsatz- and Vorsteuerabstimmung): the class-3 revenue per the books adjusted for the accruals and disposal proceeds the ledger can name against the sum of Ziffer 200 of the year's returns, and the booked Vorsteuer on 1170 + 1171 against the sum of Ziffer 400 + 405, each with its difference and a status (match, warn, or unavailable naming the periods not yet filed). Under the Saldosteuersatz the Vorsteuer half is not applicable; a workspace with no MWST method answers applicable: false. A read.
| Name | Required | Description | Default |
|---|---|---|---|
| year | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description adds substantial behavioral context: the exact computations, the source fields (Ziffer 200, 400+405, accounts 1170/1171), the status values ('match', 'warn', 'unavailable'), and the condition when it returns 'applicable: false' (Saldosteuersatz or no MWST method). It also explicitly ends with 'A read.' confirming safety.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence adds substantive detail—legal references, computation logic, status meanings, and method exceptions. It is front-loaded with the main purpose and flows logically. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of the reconciliation and the absence of an output schema, the description covers the inputs, computation, statuses, and edge cases. It does not explicitly describe the response JSON structure, but that is partially mitigated by the level of detail about what is returned. It is adequate for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must explain the parameters. It mentions 'calendar year' and 'workspace' implicitly, but does not explicitly state that workspaceId is the workspace identifier and year is the calendar year string. The mapping is inferable but not spelled out, leaving room for ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states what the tool returns: the two annual MWST reconciliations as figures, with differences and statuses. It distinguishes from siblings by specifying it covers the annual reconciliation (Art. 128 MWSTV) rather than periodic returns or previews, and mentions the Saldosteuersatz exception.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for annual reconciliation tasks and is clearly read-only. It does not explicitly name alternatives or state when not to use it, but the domain-specific detail makes it unambiguous which VAT reconciliation it addresses.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_code_deactivateA
Archive a tax code (never deletes; a posted line always resolves its code).
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It explicitly states the tool 'never deletes' and explains the reason ('a posted line always resolves its code'), which is valuable behavioral context. It does not mention other consequences like impact on future transactions, but for a simple action, this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One concise sentence that states the action, the resource, and the key behavioral guarantee. Extremely front-loaded with no waste, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with two simple self-explanatory parameters and no output schema, the description is nearly complete. It explains the action and the non-destructive behavior. It could mention whether reactivation is possible, but the purpose is clear enough for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. The description identifies the target (tax code) and implies the action, but it does not explain parameters. However, the parameters 'workspaceId' and 'code' are self-explanatory given the tool's purpose, and the description clarifies the effect. The lack of explicit parameter details is mild given the simple schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb 'Archive' and resource 'tax code', clearly distinguishing it from 'vat_code_reactivate' and 'vat_code_upsert' in the sibling list. It also clarifies that it does not delete, which is a key differentiator.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for archiving tax codes, which is clear enough for when to use it, but it does not explicitly state when not to use it or mention alternatives like 'vat_code_reactivate' or 'vat_code_upsert' for different scenarios. No exclusions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_code_reactivateA
Reactivate an archived tax code, so it appears again on new documents. The exact mirror of vat_code_deactivate; posted lines never depended on the flag.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden. It discloses a useful behavioral nuance ('posted lines never depended on the flag'), but omits other transparency aspects like error conditions, idempotency, permissions, or what happens if the code is already active. This is a partial disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no fluff. The core action and effect are front-loaded, and the mirror note adds context without redundancy. It is appropriately sized for a simple mutation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the tool is simple, the description lacks parameter documentation and error conditions. It covers the purpose and a key behavioral note, but without annotations or an output schema, the agent has limited guidance on usage prerequisites and return behavior. Adequate but with clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention either parameter (workspaceId or code). The names are self-explanatory, but no format, constraints, or usage details are provided, leaving the agent to guess. The description adds no parameter-level value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (reactivate), the resource (archived tax code), and the effect (appears again on new documents). It also explicitly distinguishes itself from the sibling vat_code_deactivate by calling itself the exact mirror, making the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (as the inverse of vat_code_deactivate) but does not explicitly state exclusions or conditions, such as only being applicable to archived codes or when a reactivation is desired. Still, the mirror reference provides clear contextual guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_codesBRead-only
List tax codes (active only by default).
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| includeArchived | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already declares this is a safe read operation, lowering the bar. The description adds one useful behavioral detail — that archived codes are excluded by default and can presumably be included via the includeArchived parameter — which goes beyond what the annotation states. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with zero filler. The core operation comes first and the default-filter qualifier is appended efficiently. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with a readOnlyHint annotation and no output schema, the description covers the basic purpose and default filtering behavior. However, with 0% parameter documentation in the schema, the lack of any explanation of workspaceId and includeArchived leaves a noticeable gap for an agent preparing a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, but it only touches the archived-filter concept indirectly. It never mentions workspaceId (required) or includeArchived by name, and it does not explain what 'active' means in relation to the parameters. The hint about active-only is helpful but insufficient for a required parameter and a boolean flag.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('List') and resource ('tax codes'), so an agent knows exactly what operation this performs. The parenthetical '(active only by default)' adds a behavioral qualifier that helps distinguish it from sibling tools like vat_code_upsert or vat_code_deactivate, though it does not name an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus the many VAT-related siblings (vat_codes, vat_config, vat_code_upsert, vat_configure, etc.). The use case is only implied by the verb 'List', with no exclusions or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_code_upsertC
Add or edit a tax code at the single tax enumeration point.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | ||
| kind | Yes | ||
| label | No | ||
| rateBp | Yes | ||
| formLine | Yes | ||
| validFrom | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It states 'Add or edit' which implies mutation, but does not disclose side effects, idempotency behavior (despite an idempotencyKey parameter), or any constraints. Minimal behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence, which is efficient, but it lacks structure and fails to front-load any additional details. It is appropriately short but overly minimal.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 8 parameters, no schema descriptions, no output schema, and no annotations, this description is severely inadequate. It does not mention required fields, parameter meaning, or expected behavior, making it impossible to call correctly without external knowledge.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the description provides no parameter explanations. It does not explain what 'code', 'kind', 'rateBp', 'formLine', or other fields mean, leaving the agent without necessary semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Add or edit' (upsert) and the resource 'tax code', plus the location 'single tax enumeration point', which distinguishes it from siblings like vat_code_deactivate and vat_code_reactivate that deal with status changes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'at the single tax enumeration point' provides some context, but there is no explicit guidance on when to use this tool versus related VAT tools (vat_codes, vat_code_deactivate, vat_code_reactivate). It is implied but not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_configARead-only
Read the current VAT configuration, including the Saldosteuersätze in force today, each Tätigkeit with the Ertragskonten mapped to it, and the elected MWSTV Art. 88 Abs. 6 declaration basis.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint already disclosing that this is a safe read, the description adds useful behavioral context: results are current as of today, the response includes mappings and the elected declaration basis. It does not describe edge cases such as missing configuration, but the added scope detail is meaningful and consistent with the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, well-structured sentence front-loads the verb and resource and then appends specific, non-redundant details. There is no filler, and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only tool with no output schema, the description covers the main return contract by listing the three key components. Minor gaps remain, such as behavior when no VAT configuration exists or explicit confirmation that workspaceId is required, but the description is complete enough for correct invocation in most cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one required parameter, workspaceId, with 0% description coverage and no enums or value hints. The description never mentions workspaceId or what it selects, so it adds no parameter semantics beyond the parameter's self-explanatory name and type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the exact operation (read) and resource (current VAT configuration), then names three concrete content areas: Saldosteuersätze in force today, Tätigkeit-to-Ertragskonten mappings, and the elected MWSTV Art. 88 Abs. 6 declaration basis. This clearly separates it from write/setup siblings like vat_configure and narrower reads like vat_saldo_declaration_basis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use guidance or alternatives are provided. The read-only nature is implied by the verb and the readOnlyHint annotation, but the description does not tell an agent when to choose this tool over related vat_* siblings, nor when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_configureA
Configure VAT: method, timing, registration, VAT number, and the ESTV Bewilligung (asOf picks the statutory rate era). A filer holding several Saldosteuersätze maps each Tätigkeit to its Ertragskonten with saldoActivities, which is how MWSTV Art. 84 Abs. 3 books the Erträge separately per rate. Changing an approval that already governs posted turnover must say which it is: saldoGrant for a new ESTV Bewilligung from a given day, saldoCorrection to rewrite the open one.
| Name | Required | Description | Default |
|---|---|---|---|
| asOf | No | ||
| method | Yes | ||
| timing | Yes | ||
| vatNumber | No | ||
| registered | Yes | ||
| saldoGrant | No | ||
| saldoRates | No | ||
| workspaceId | Yes | ||
| methodChange | No | ||
| idempotencyKey | Yes | ||
| saldoActivities | No | ||
| saldoCorrection | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it does add meaningful behavior: asOf selects the statutory rate era, saldoActivities maps Tätigkeiten to Ertragskonten, and saldoCorrection rewrites the open approval. It does not disclose auth needs, reversibility, or output behavior, but the described distinctions go well beyond the bare schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences long, front-loaded with a concise summary, and every sentence adds substantive information rather than repeating the tool name or schema. The legal citation is dense but directly explains why saldoActivities is structured the way it is.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity—12 parameters, nested objects, no annotations, and no output schema—the description provides strong domain context for the saldo-related sub-features but is incomplete for a complete call. It omits the semantics of several required and optional parameters and does not mention what happens after a successful configuration.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining parameters. It does add semantics for asOf, saldoActivities, saldoGrant, and saldoCorrection, and it lists method, timing, registration, and vatNumber. However, it leaves several parameters unexplained, notably methodChange, saldoRates, workspaceId, and idempotencyKey, and it gives no value constraints for method or timing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as configuring VAT and enumerates the key facets: method, timing, registration, VAT number, and the ESTV Bewilligung. This is a specific verb-plus-resource statement, though it does not explicitly differentiate itself from sibling tools like vat_config or set_vat_method.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides useful conditional guidance, such as when saldoGrant versus saldoCorrection should be used for an approval that governs posted turnover. However, it does not state when to prefer this tool over alternatives like vat_config or set_vat_method, so usage context is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_export_ech0217ARead-only
Export the MWST-Abrechnung for a period as an eCH-0217 v2.0.0 XML file for upload to the ESTV ePortal. Produces a file and transmits nothing: there is no ESTV submission API.
| Name | Required | Description | Default |
|---|---|---|---|
| periodEnd | Yes | Last day of the reported period, ISO YYYY-MM-DD (inclusive). | |
| periodStart | Yes | First day of the reported period, ISO YYYY-MM-DD (inclusive). | |
| workspaceId | Yes | ||
| typeOfSubmission | No | eCH-0217 typeOfSubmission: 1 Ersteinreichung (default), 2 Korrekturabrechnung, 3 Jahresabstimmung. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description adds valuable context: it produces a file and transmits nothing, explicitly ruling out an ESTV submission API. This prevents an agent from assuming network transmission. However, it does not elaborate on file delivery (e.g., return value, storage location) or any prerequisites, leaving some behavioral details uncovered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded: the first sentence states the core function and format, and the second clarifies a critical non-transmission detail. No wasted words, and the most important information appears immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the tool is relatively simple (read-only export), the lack of an output schema means the agent does not know what the tool returns (e.g., XML content, file path, or download URL). The description does not mention return format or any prerequisites (e.g., VAT data prepared for the period). Given the absence of output schema, more detail would be needed for full completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides descriptions for periodStart, periodEnd, and typeOfSubmission, covering 75% of parameters. The description adds no additional parameter-specific meaning, leaving workspaceId undocumented. Given the high schema coverage, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool exports the MWST-Abrechnung for a period as an eCH-0217 v2.0.0 XML file, specifying the verb, resource, format, and target (ESTV ePortal). It also distinguishes itself from potential siblings by explicitly noting it does not transmit, making its purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clarifies that the tool only produces a file and does not submit to ESTV, which hints at its non-transmission role. However, it does not explicitly name alternatives like export_vat or vat_return, nor does it state when to prefer this tool over them. The guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_mark_filedA
Record that a VAT period was filed with the ESTV, applying the A03 hard lock to its months. Transmits nothing. CONSEQUENCE: Marks the VAT period as filed with the ESTV; a filed period is closed to correction runs.
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | Period label: YYYY-Qn (effektiv) or YYYY-Hn (Saldo). | |
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses a significant consequence: 'applying the A03 hard lock to its months' and 'a filed period is closed to correction runs.' It also explicitly states 'Transmits nothing,' preventing a false assumption. It omits details like idempotency effects or reversibility, but the core behavioral impact is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, each with a distinct purpose: action+lock, non-transmission, and consequence. The most important information is front-loaded, and there is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-parameter tool with no annotations and no output schema, the description covers the key semantics (what it does, what lock it applies, that it doesn't transmit). However, it leaves questions about prerequisites (e.g., period must be open), idempotencyKey semantics, and whether the lock can be undone. These are not essential for a basic call but are relevant for safe use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only 'period' has a description). The tool description does not explain 'workspaceId' or 'idempotencyKey' beyond their names. Given the low coverage, the description should compensate, but it only indirectly clarifies that 'period' identifies the VAT period via the lock reference. This leaves two of three parameters under-documented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Record that a VAT period was filed'), names the resource (VAT period, ESTV), and clarifies a critical distinction: 'Transmits nothing.' It also names the A03 hard lock, which makes the tool's role concrete and differentiable from actual filing tools like vat_return.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than explicit: the description says it records filing and transmits nothing, so an agent can infer it is for post-filing bookkeeping. However, it does not name alternatives (e.g., vat_return, vat_settlement_post) or state when not to use this tool, leaving some ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_periodsBRead-only
List the statutory reporting periods of a year (quarterly under effektiv, semi-annual under Saldo) and whether each is filed.
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | Calendar year, YYYY. | |
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations declare readOnlyHint=true, so the description's burden is lighter. It adds useful context about the periodicity under different VAT methods (effektiv quarterly, Saldo semi-annual) and the filing status, but does not describe edge cases or return structure. This is acceptable given the read-only annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, under 20 words, with no filler. It conveys the core purpose and key distinctions immediately, making it highly efficient and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description omits behavior when 'year' is not provided (the schema marks it optional) and does not specify the output structure. Given the existence of related VAT settlement tools, a note on how this differs or what the list contains would improve completeness. However, the core concept is clear enough for a simple read operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%: only 'year' has a description; workspaceId has none. The tool description mentions 'a year' but does not connect it to the parameter, and workspaceId is not mentioned at all. The description fails to compensate for the schema's incomplete parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and a concrete resource ('statutory reporting periods of a year') with an additional attribute ('whether each is filed'). It clearly distinguishes this tool from other VAT tools by focusing on period listing rather than calculations or submissions, even if it doesn't explicitly name siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as vat_preview or vat_return. It does not state conditions, exclusions, or alternatives, leaving the agent to infer applicability solely from the tool name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_previewARead-only
Preview a line VAT (net/tax/gross, ESTV form line, kind, deductibility) without posting.
| Name | Required | Description | Default |
|---|---|---|---|
| taxCode | No | ||
| supplyDate | No | Supply date (Leistungsdatum), ISO YYYY-MM-DD; picks the statutory rate era. | |
| amountMinor | Yes | ||
| workspaceId | Yes | ||
| amountIsGross | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations include readOnlyHint=true, which already tells the agent this is a safe read operation. The description adds the key behavioral trait that it does not post (no side effects), which aligns with the annotation. It doesn't add much beyond that—no mention of rate-limit, auth, or what happens with invalid tax codes. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, front-loaded with the verb and resource, and includes the key scoping phrase 'without posting'. Every word earns its place. It's concise and structured well.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a preview tool with readOnlyHint=true, the description is mostly complete. It tells the agent what it does and that it doesn't post. However, it doesn't explain the return value (what the preview looks like), and with no output schema, the agent might not know what to expect. Also, the parameter semantics are incomplete (e.g., how amountIsGross interacts with amountMinor). Given the tool's complexity (5 params, 2 required), a bit more detail would help.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20% (only supplyDate has a description). The description mentions 'net/tax/gross' and 'amountIsGross' is a parameter, but it doesn't explain the relationship between amountMinor and amountIsGross, nor what taxCode should be. The description adds some context (the preview includes net/tax/gross) but doesn't fully compensate for the low schema coverage. Baseline 3 is appropriate because the description adds some meaning but leaves gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Preview') and resource ('a line VAT'), and lists the key fields (net/tax/gross, ESTV form line, kind, deductibility). It clearly distinguishes itself from posting tools like vat_settlement_post by explicitly saying 'without posting'. However, it doesn't explicitly name a sibling alternative, so it doesn't fully differentiate from other preview tools like vat_settlement_preview or tax_provision_preview.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: use this when you want to preview VAT line details without posting. The phrase 'without posting' gives a clear context and exclusion. However, it doesn't explicitly state when to use this vs alternatives like vat_settlement_preview or tax_provision_preview, nor does it mention any prerequisites (e.g., workspace configuration, tax codes).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_returnBRead-only
Compute the MWST-Abrechnung for a period: ESTV form lines with base/tax, the payable or credit, and the drill-down entry ids.
| Name | Required | Description | Default |
|---|---|---|---|
| periodEnd | Yes | Last day of the reported period, ISO YYYY-MM-DD (inclusive). | |
| periodStart | Yes | First day of the reported period, ISO YYYY-MM-DD (inclusive). | |
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the safety profile is covered. The description's 'Compute' aligns with the read-only annotation and adds no side-effect context, which is acceptable given the annotation, but it contributes little beyond the annotation itself.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that efficiently conveys the verb, resource, and output components with no filler. Every phrase earns its place and the key action is stated first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description gives a useful sketch of the return output, which is helpful given the absence of an output schema. However, it omits any guidance distinguishing it among the many VAT-related siblings and any calculation prerequisites, leaving the tool's place in the workflow incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67%: periodStart and periodEnd are already described in the schema, leaving workspaceId undocumented with no compensating detail. The description repeats the 'period' concept without adding meaning beyond the schema, so it doesn't help the agent understand workspaceId or any other parameter better.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Compute'), a resource ('MWST-Abrechnung'), and a period scope, then lists the output components (ESTV form lines, payable/credit, drill-down entry ids). This is clear and precise, though it does not explicitly name sibling tools for differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus closely related siblings like vat_preview, vat_settlement_preview, or vat_export_ech0217. The description implies it computes the official VAT return but doesn't state exclusions, prerequisites, or alternative routing, leaving the agent to infer selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_saldo_declaration_basisA
Elect, or withdraw, the MWSTV Art. 88 Abs. 6 simplification for one Steuerperiode: highest_rate declares the whole taxable turnover at the highest approved Saldosteuersatz on one Ziffer, per_activity restores the Abs. 1 split per Tätigkeit. It is voluntary and usually raises the tax due, so the engine never applies it on its own.
| Name | Required | Description | Default |
|---|---|---|---|
| basis | Yes | ||
| taxPeriod | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It does add useful context: the operation is voluntary, typically increases tax due, and will never be applied automatically by the engine. But it does not disclose other behavioral aspects such as reversibility, effects on future periods, or what happens to previously applied simplifications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a tight two-sentence structure that immediately states the action, explains the two modes, and then covers the key usage warning. No filler or redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is a state-changing VAT election, and the description covers its core semantics, the two basis modes, and the tax impact. It does not mention eligibility (e.g., whether the simplification is available in the first place) or what idempotencyKey is for, but those could be expected from the context or other tools. Overall it is sufficiently complete for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% coverage, so the description must compensate. It explains the `basis` parameter values ('highest_rate' and 'per_activity') and ties the operation to a single tax period (`Steuerperiode`), clarifying `taxPeriod`. The other parameters (`workspaceId`, `idempotencyKey`) are not detailed, but the two most consequential ones are addressed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Elect, or withdraw') on a clearly identified resource ('MWSTV Art. 88 Abs. 6 simplification for one Steuerperiode'). It also contrasts the two modes ('highest_rate' vs 'per_activity'), which uniquely distinguishes this tool from any VAT-related sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context on when to use the tool: 'It is voluntary and usually raises the tax due, so the engine never applies it on its own.' This signals that a user must explicitly request this change and should be aware of the tax impact. However, it does not name alternative tools (e.g., vat_saldo_eligibility) or explicitly state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_saldo_eligibilityCRead-only
The MWSTG Art. 37 Abs. 1 Saldo eligibility limits in force today (both halves of the cumulative test, era-scoped, boundary 1.1.2024) beside the measured steuerbarer Umsatz of one calendar year (Ziffer 299, computed by the same path as the Abrechnung). Never a verdict: eligibility turns on EXPECTED turnover (ESTV practice, MWST-Info 12) and on a tax-due half that needs a rate the ESTV has not granted yet, so this read compares and stops.
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description doesn't need to repeat that. It adds useful behavioral context: that this tool 'compares and stops' and does not issue a verdict, and that eligibility depends on expected turnover and an ungranted rate. This goes beyond the read-only hint and clarifies the tool's scope and limitations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is verbose and dense, packed with legal references and technical terms (MWSTG Art. 37 Abs. 1, Ziffer 299, era-scoped, boundary 1.1.2024). While it front-loads the main purpose, the overall text is overly complex and not concise. An agent would need significant parsing to extract the core message, and the style hinders quick understanding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description should explain what the tool returns. It states it provides the limits and the measured turnover, and that it 'compares and stops', which gives some idea. However, it does not specify the output structure, does not explain the parameters, and omits details about how the comparison is presented. Given the tool's complexity and the absence of output schema, this is incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% – the description does not mention the 'year' or 'workspaceId' parameters at all. It refers to 'one calendar year' but never explicitly ties that to the 'year' input, and it doesn't explain the required workspaceId or the expected format for year. Since the description must compensate for the missing schema descriptions, this is a severe gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific purpose: comparing MWSTG Art. 37 Abs. 1 Saldo eligibility limits with the measured steuerbarer Umsatz. It is domain-specific and distinguishes itself from other vat tools (e.g., vat_preview, vat_return) by focusing on eligibility comparison, not on filing or calculation. However, the heavy jargon makes it less accessible without domain knowledge, so it's not a full 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (when needing the comparison of limits and measured turnover) and explicitly warns 'Never a verdict', which is a clear exclusion. But it does not name alternative tools or specify conditions for choosing this over siblings like vat_saldo_generations or vat_settlement_preview. The guidance is implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_saldo_generationsARead-only
The Bewilligungsverlauf: every ESTV Saldosteuersatz approval this workspace has recorded, oldest first, with the days each governed and the Tätigkeiten and Ertragskonten it carried. Under Saldo no rate is stamped on a journal line, so this history is the only evidence of what a filed period was computed with.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, and the description does not contradict this. It adds value beyond the annotation by specifying the output contents (days, activities, accounts), ordering (oldest first), and the business reason. No side effects or auth details, but read-only nature is covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no fluff. The main purpose is front-loaded, then details and rationale. It is appropriately sized for a simple read-only list tool, though it could be slightly more compact without losing value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given one parameter, no output schema, and read-only annotation, the description conveys the return contents (list of approvals with specific fields) and the ordering. It does not mention pagination or limits, but for a simple list tool this is acceptable. The business context is also explained, making it complete for agent usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (workspaceId) with no description coverage. The description implies workspace context ('this workspace') but does not explicitly explain the parameter. Since it's a trivial identifier, the schema's requirement is self-evident, but the description does not add explicit semantics. Baseline 3 is appropriate given the minimal schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists all ESTV Saldosteuersatz approvals (Bewilligungsverlauf) for the workspace, oldest first, with specific details (days, Tätigkeiten, Ertragskonten). It names a specific verb (list) and resource, and the rationale (only evidence of computed rate) distinguishes it from related VAT tools like vat_saldo_eligibility.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides strong context: under Saldo method no rate is stamped on journal lines, so this history is the only evidence of what a filed period used. This implicitly tells when to use it (when needing to verify a filed period's rate) but does not explicitly name alternatives or exclusion conditions. Clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_seed_defaultsB
Seed the default Swiss tax-code set (idempotent).
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries full burden. It discloses idempotency, which is valuable, but does not state whether it creates missing codes only or resets existing ones, nor any permission or side-effect details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with the action and resource front-loaded. It is appropriately sized, though it omits details that would make it more useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool, the description gives minimal context. It doesn't explain what the default Swiss tax-code set contains, what happens to existing codes, or what the expected outcome is, leaving an agent to infer the operation's effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not mention the workspaceId parameter. Although the name is self-explanatory, the description adds no meaning about how it scopes the seeding operation or whether a specific workspace state is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('seed') and a specific resource ('default Swiss tax-code set'), and notes idempotency. It is clearly distinguishable from sibling VAT tools like vat_configure or vat_code_upsert, which operate on configuration or individual codes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to call this tool, what prerequisites exist, or how it relates to alternatives. There is no mention of whether it should be run during onboarding or before using other VAT features.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_settlement_listARead-only
List the MWST-Saldierungen on record, newest period first, posted and reversed, with the moved figures and the entry ids. year (YYYY) narrows to one Steuerperiode. A read.
| Name | Required | Description | Default |
|---|---|---|---|
| year | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark readOnlyHint=true, and the description reinforces this with 'A read.' It adds valuable behavioral context beyond the annotation: results are sorted newest first, reversed settlements are included, and the response contains moved figures and entry IDs. It does not mention pagination or rate limits, but the read-only safety profile is well covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single information-dense sentence plus the short clarification 'A read.' It front-loads the core purpose and includes ordering, inclusivity, fields, and filtering without wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does well by naming the returned data (moved figures, entry IDs) and specifying ordering and inclusion of reversed entries. The required workspaceId is not described, and pagination is absent, but the tool is a simple list and the read-only annotation reduces the burden. Overall, it is complete enough for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain the year parameter with format and meaning ('year (YYYY) narrows to one Steuerperiode'), but it says nothing about the required workspaceId parameter beyond what the schema's name implies. This leaves a gap in parameter-level guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('List'), a specific resource ('MWST-Saldierungen'), and adds concrete detail: newest period first, posted and reversed entries, moved figures, and entry IDs. This clearly distinguishes it from related settlement tools like vat_settlement_post or vat_settlement_reverse.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it by saying 'A read' and by describing list behavior, and it explains the optional year filter. However, it does not explicitly name sibling tools or state when not to use it versus vat_settlement_preview/post/reverse. Usage context is present but exclusions are left implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_settlement_postA
Book the MWST-Saldierung the preview showed: transfer the filed period's balances on 2200, 1170 and 1171 to 2201 (MWST-Abrechnungskonto), dated the period end, source='vat_settlement', so the three tax accounts read zero and 2201 carries what the ESTV is owed. Admitted inside the filed, hard-locked period under three enforced conditions (no VAT trace, only the tax accounts, never into a year-close seal); the filed return and the Abstimmung are unchanged by it. Refuses period_not_filed before the filing, nothing_to_settle on an empty period, already_posted while a settlement of the period stands, period_locked with reason year_close on a sealed year. Idempotent on idempotencyKey for as long as the settlement that key booked still stands; a key whose settlement was reversed refuses already_reversed_key (post again under a new key, which books a new settlement). The correction is vat_settlement_reverse. CONSEQUENCE: Transfers the filed period's VAT balances from 2200, 1170 and 1171 to 2201 inside the filed period; the filed return does not change, and the only correction is a reversing entry.
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description thoroughly discloses behavioral traits: it forces source='vat_settlement', operates only under three conditions (no VAT trace, only tax accounts, never into year-close seal), is idempotent on idempotencyKey, and lists refusal reasons and reversibility behavior. Since no annotations are provided, the description carries the full burden, and it does so richly, covering idempotency, error conditions, and that the filed return remains unchanged. The only slight gap is not explicitly stating the effects on the input period's balance outward, but it is covered well beyond minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense with valuable information but is somewhat long and laborious, covering many edge cases (all refusal reasons, idempotency details, consequences) in a single paragraph without clear structure. It is front-loaded with the primary action, which is good, but the middle section enumerates many technical details that could be condensed or formatted. It earns its place but could be more scannable; still, it is not rambling or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (VAT settlement posting with multiple constraints, idempotency, and error codes) and the lack of annotations and output schema, the description is nearly complete. It explains prerequisites (filed period, preview done), conditions, failure modes, idempotency, and consequences (accounts zeroed, 2201 carries debt). The only minor absence is a note about what happens to VAT trace or audit trail, but the mention of conditions and the reversing entry suffices. It covers all essential aspects for an agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and parameters are just strings without descriptions, so the description must add semantic meaning. It does: it references the 'period' as the filed period, explains 'idempotencyKey' behavior (idempotent as long as settlement stands; refused with 'already_reversed_key' if reversed), and implies 'workspaceId' is the context. It also mentions the need for prior preview ('the preview showed') and the period end date. This exceeds the schema by providing domain-specific semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Book the MWST-Saldierung the preview showed') and the resource (transfer balances from accounts 2200, 1170, 1171 to 2201), with specific quoting of source and dates. It distinguishes itself from the sibling 'vat_settlement_preview' by focusing on the actual posting step, and from 'vat_settlement_reverse' as the correction tool. The verb 'book' and detailed account numbers make the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use it ('Admitted inside the filed, hard-locked period under three enforced conditions') and when not to use it (e.g., refuses before filing with 'period_not_filed', or if period is locked with 'year_close'). It also names the alternative 'vat_settlement_reverse' as the correction. These conditions and error cases give clear guidance for an agent to decide when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_settlement_previewARead-only
Preview the MWST-Saldierung of one filed period (period is the A07 label, 2026-Q3 or 2026-H1): the booked movement on 2200 (Umsatzsteuer), 1170 and 1171 (Vorsteuer) over the period beside the return's Ziffer 399 and 400 + 405 with the differences, and the exact lines the post would book (Dr 2200 / Cr 2201, Dr 2201 / Cr 1170, Dr 2201 / Cr 1171; under the Saldosteuersatz the flat-rate tax due against 3809). Settlement entries and their reversals are excluded from the read, so a settled period previews as settled, never twice. A read: nothing is posted.
| Name | Required | Description | Default |
|---|---|---|---|
| period | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description reveals the important behavioral nuance that settlement entries and their reversals are excluded, so a settled period previews as settled, never twice. It also specifies exactly what bookkeeping lines would be posted and reiterates that nothing is posted. This goes well beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is detailed yet efficiently structured: it opens with the core purpose, defines the period format, explains the preview contents, calls out the exclusion of settlement entries, and ends with a read-only assurance. Each sentence adds necessary information, and there is no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema and only a readOnlyHint annotation, the description fully carries the burden of explaining the tool's behavior. It clearly describes the inputs, the detailed contents of the preview (accounts, Ziffer numbers, differences, booking lines), special edge-case handling, and confirms no post is made. This is sufficient for an agent to understand the operation and its result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, leaving both parameters undocumented. The description compensates by explaining the period parameter format ('A07 label, 2026-Q3 or 2026-H1') and its role as a filed period. workspaceId is not described, but it is a recurring global parameter across the sibling toolset, so the description meaningfully clarifies the more ambiguous parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb and resource: 'Preview the MWST-Saldierung of one filed period.' It details what the preview includes (movements on accounts 2200, 1170, 1171; Ziffer 399 and 400+405; differences; and booking lines) and clearly states 'A read: nothing is posted,' distinguishing it from posting tools like vat_settlement_post.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the usage context clear by labeling it a 'Preview' and emphasizing read-only behavior. It implies this is for analysis before any posting, but does not explicitly name alternative tools or state conditions such as 'use vat_settlement_post for actual posting.' The guidance is strong but not fully explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vat_settlement_reverseA
Reverse a posted MWST-Saldierung with a reversing entry dated the period end, so 2200, 1170, 1171 and 2201 net to zero again inside the settled period; the settlement row reads reversed and keeps its history, and a later vat_settlement_post for the period books a new settlement. This is the ONLY way to reverse a settlement entry: the generic reverse_entry refuses it with owned_by, because the settlement row must move with the mirror. Refuses not_found and already_reversed. Idempotent on idempotencyKey. CONSEQUENCE: Reverses the VAT settlement with a mirror entry dated the period end; the settlement stays on record as reversed.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes | ||
| settlementId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure, and it delivers: it explains the reversing entry is dated period end, the settlement row is marked reversed while keeping history, the tool is idempotent on idempotencyKey, and it refuses with not_found and already_reversed. These are meaningful, non-obvious behaviors that an agent needs to call and interpret the tool correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the primary action, then gives routing, error, and idempotency details. However, the final CONSEQUENCE sentence largely repeats the opening sentence ('Reverses the VAT settlement with a mirror entry dated the period end; the settlement stays on record as reversed'), adding redundancy. Otherwise, every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema and no annotations, this description covers the essential context: what the tool does, which accounts net out, what happens to the settlement row, when it can fail, how idempotency works, and why the generic reverse_entry cannot be used. Nothing critical is missing for an agent to invoke and understand the result of this reverse operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains idempotencyKey by stating 'Idempotent on idempotencyKey' and implies settlementId is the identifier of the settlement being reversed, but it never explicitly names or defines workspaceId and settlementId. The parameter names are fairly self-evident, but the description leaves some burden on the schema's bare property names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource: 'Reverse a posted MWST-Saldierung' and immediately explains the operational effect on specific accounts (2200, 1170, 1171, 2201). It distinguishes itself from the generic reverse_entry by explicitly stating it is the only way and why, so an agent can differentiate this tool from siblings without inspecting schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool versus the generic reverse_entry: 'This is the ONLY way to reverse a settlement entry: the generic reverse_entry refuses it with owned_by.' It also gives post-conditions like 'a later vat_settlement_post for the period books a new settlement,' which helps an agent choose the right sequence. No ambiguity remains about when this tool applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vendor_portal_grantA
Gib einem Lieferanten Portal-Zugang frei (US-F03.1): mint a scoped, expiring vendor grant over F02's shared grant engine (kind=vendor, no second token path). The contact must be a SUPPLIER (party_role vendor|both), else contact_not_found (the same opaque code a missing contact returns, no oracle). scopes default to the vendor set [{kind:"pos.read"},{kind:"remittance.read"}]. The token is CSPRNG >=256-bit; only its SHA-256 hash is stored and the one-time link is returned once (tokenOnce/localLink), never re-derivable. expiresAt in the past is refused (expiry_in_past); longer than the 90-day max is CLAMPED (clamped:true). Posts nothing (OP4: mints the local artifact and stops, hosted:false/reason:cloud_tier). idempotencyKey replays the original grant.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| scopes | No | ||
| contactId | Yes | ||
| expiresAt | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description fully carries the behavioral disclosure burden and does so thoroughly: token generation is CSPRNG >=256-bit, only the SHA-256 hash is stored, the one-time link is returned only once and is not re-derivable, expired expiresAt is refused, long expiry is clamped, and the operation posts nothing. This goes far beyond minimal disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-dense, with every clause contributing to correct invocation: scoping, conditions, defaults, security, expiry handling, idempotency, and side-effect behavior. It is somewhat dense with internal codes and semicolon-packed clauses, but the structure is logical and front-loaded with the core action.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has no output schema and no annotations, the description supplies substantial context: error codes, security model, clamping behavior, post-nothing side-effect, and the one-time-link return path. It does not fully describe the output object or every parameter, but it provides enough for an agent to call this correctly and interpret the main response fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds real semantics for contactId (supplier role requirement), expiresAt (past refusal, 90-day clamp), scopes (default vendor set), and idempotencyKey (replays original grant). workspaceId and actor are not individually explained, leaving a small gap, but the description handles the most decision-critical parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: "Gib einem Lieferanten Portal-Zugang frei" — mint a scoped, expiring vendor grant over F02's shared grant engine (kind=vendor, no second token path)." It clearly distinguishes this tool from generic portal grant paths and from sibling vendor_portal_revoke/grants_list by specifying the vendor-specific grant kind.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete preconditions: the contact must be a SUPPLIER (party_role vendor|both), otherwise contact_not_found is returned. It also defines behavioral contracts like default scopes, expiry clamping, and idempotencyKey replay, so an agent knows when and how to call it. It does not, however, explicitly name alternative tools for the non-vendor or generic portal-grant case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vendor_portal_grants_listARead-only
Liste die Lieferanten-Portalfreigaben (P5): every vendor grant, optionally per contact and per status (draft/active/revoked/expired), each with its hosted:false truth. Reuses F02's grant read model filtered to kind=vendor; savedViewId applies a G00 saved view over the portal_grant entity kind.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | ||
| contactId | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true already declaring the safety profile, the description adds useful scoping context: the hosted:false truth, the F02 grant read model reuse, and the G00 saved-view application. But it omits behavioral details an agent might want, such as pagination behavior, result ordering, or whether the read model excludes certain grants by default.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences that front-load the core purpose and filter capabilities before the implementation notes. Every clause earns its place, though the internal references (F02, G00) introduce jargon that a non-expert agent may need to resolve.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with readOnlyHint and no output schema, the description is mostly adequate: filters, status values, and scoping are covered. However, with no output schema, return shape and pagination are not described at all, and the expected response contents are only vaguely implied by 'each with its hosted:false truth.'
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden — and it compensates well: it documents that status accepts draft/active/revoked/expired values (crucial, since the schema has no enums), explains contactId filters per contact, and explains that savedViewId applies a G00 saved view over the portal_grant entity kind. Only workspaceId is left unexplained, which is a common contextual parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb (Liste/list) and resource (Lieferanten-Portalfreigaben / vendor grants), and clarifies scope: every vendor grant with optional contact/status filtering and hosted:false semantics. It partially distinguishes itself from siblings like portal_grant_list via 'filtered to kind=vendor' and 'hosted:false truth', though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied: the description states the available filters (per contact, per status) and the saved-view behavior, which tells an agent what this call is for. However, there is no explicit when-to-use vs. alternatives guidance — it never says 'use X instead for hosted grants' or names portal_grant_list as the sibling covering hosted grants.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vendor_portal_posARead-only
Liste die offenen Bestellungen eines Lieferanten (US-F03.2, P5): a scoped read model over D02 purchase_orders + po_lines, filtered to the supplier's own contact, status IN (sent, received) only (draft/closed/cancelled are D02 internal state, never exposed), and workspace_id (H-TENANT). Pass grantToken for the supplier/agent path (the grant's own contact_id is the fence; an expired/revoked/foreign token returns grant_invalid, no oracle; the grant must carry the pos.read scope) OR contactId for the operator "Sichtbar für Lieferant" preview. Reads only, mutates nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| contactId | No | ||
| grantToken | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=true already set, the description adds meaningful context: it explains token failure behavior (expired/revoked/foreign token returns grant_invalid, no oracle), the required pos.read scope, and that certain statuses are never exposed. This goes beyond the annotation and clarifies the read-only nature without contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient, front-loading the purpose and then packing in scope, filters, auth modes, and safety in roughly three sentences. Some jargon (US-F03.2, D02, H-TENANT) adds precision, though it may slightly reduce readability for newcomers; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with three parameters and no output schema, the description covers the essential context: data source, filtering rules, authentication paths, error behavior, and side-effect safety. It doesn't describe the return shape or pagination, but the lack of an output schema makes that less critical; overall it gives enough for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does: it explains workspace_id as the H-TENANT filter, grantToken as the supplier/agent auth path with the scope requirement, and contactId as the operator preview path. It does not spell out types or mutual exclusivity, but it conveys the semantic role of each parameter far better than the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Liste die offenen Bestellungen eines Lieferanten' (list a supplier's open orders). It further distinguishes this from sibling tools by defining it as a 'scoped read model' over D02 purchase_orders + po_lines with explicit status and workspace filters, so an agent can tell it apart from other vendor_portal_* and PO tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit routing guidance: 'Pass grantToken for the supplier/agent path ... OR contactId for the operator preview.' It also clarifies which statuses are excluded ('draft/closed/cancelled are D02 internal state, never exposed'), telling the agent when not to rely on this tool. It does not name specific alternative tools for those cases, so a small gap remains.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vendor_portal_remittance_createA
Erstelle ein Zahlungsavis (US-F03.3): SNAPSHOT the A14 payment and its A17 vendor-bill allocations into a remittance_advice + per-bill lines (integer-Rappen values copied verbatim, per-line H-FX txn+base+rate), and file a rendered advice artifact into E00 (artifact_document_id). POSTS NOTHING: A14 already posted the payment (P3 by having no posting at all). The paymentId must be an OUTGOING supplier settlement with >=1 vendor-bill allocation, else payment_not_found (opaque, no cross-supplier oracle). idempotencyKey (scoped by paymentId) replays the same advice; without a key a re-file supersedes the current advice (supersedes_id), never a silent overwrite. Artifact-and-stop (OP4: transmitted:false/reason:cloud_tier).
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| paymentId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full transparency burden and does so exceptionally. It discloses that the tool POSTS NOTHING, files an artifact, supports idempotent replay via idempotencyKey, supersedes prior advice instead of silently overwriting, and sets transmitted:false/reason:cloud_tier. It also reveals the opaque payment_not_found error condition.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-packed, with every clause earning its place, including the no-posting guarantee, idempotency behavior, and artifact stop condition. It is front-loaded with the core purpose, though the heavy use of cryptic identifiers like A14, A17, P3, OP4, and E00 makes it less immediately readable than it could be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex, side-effectful tool with no annotations and no output schema, the description covers purpose, preconditions, error semantics, side effects, idempotency, and file/artifact behavior. The main omissions are the exact response shape beyond artifact_document_id/supersedes_id and any explicit mention of workspaceId semantics, but the overall guidance is strong.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Despite 0% schema description coverage, the description gives rich semantics for paymentId and idempotencyKey: paymentId must be an outgoing supplier settlement with allocations, and idempotencyKey is scoped by paymentId with replay behavior. workspaceId and actor are not described, but their names are self-explanatory and the required-schema already marks workspaceId as required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource: creating a remittance advice from an A14 payment and its A17 vendor-bill allocations, then filing a rendered artifact into E00. It is clearly distinct from sibling tools like vendor_portal_remittances or vendor_portal_grant by focusing on the create-and-file action with specific source data.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit preconditions for correct use: paymentId must be an outgoing supplier settlement with at least one vendor-bill allocation, and the tool posts nothing because the A14 payment is already posted. It does not explicitly name alternative sibling tools or exclusion cases, but the operational context is clear enough for an agent to decide when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vendor_portal_remittancesARead-only
Liste die Zahlungsavise eines Lieferanten (US-F03.3 portal read, P5): the remittance advices scoped to the supplier's own contact, same fences as vendor_portal_pos (grantToken requires the remittance.read scope, or contactId for the operator preview). Each advice carries its per-bill lines with txn+base+rate. savedViewId applies a G00 saved view over remittance_advice (its stored presets never widen the contact fence). Reads only.
| Name | Required | Description | Default |
|---|---|---|---|
| contactId | No | ||
| grantToken | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description adds meaningful behavioral detail: it is explicitly read-only, it never widens the contact fence even when a saved view is applied, and each advice carries per-bill lines with txn+base+rate. This gives an agent important expectations about side effects and data shape beyond what the annotation alone provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but not bloated; every sentence adds a distinct fact about scope, auth, output, or saved-view behavior. The mixed German/English phrasing and the packed parentheticals slightly reduce readability, but the purpose is front-loaded and no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description compensates well by describing the returned remittance advice lines (per-bill lines with txn+base+rate) and the auth/scoping constraints. It does not mention pagination, ordering, or the semantics of the required workspaceId, but the essential call-decision information is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the burden, and it does so for three of the four parameters: grantToken requires remittance.read, contactId is used for operator preview, and savedViewId applies a G00 saved view without widening the fence. The required workspaceId is not explained, which is a minor gap for a required parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear specific verb and resource: 'Liste die Zahlungsavise eines Lieferanten' (list the remittance advices of a supplier). It also differentiates the scope from sibling tools by stating the data is scoped to the supplier's own contact and by referencing vendor_portal_pos as the same fence, so an agent can distinguish this list operation from related portal tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when the tool is appropriate: it is a portal read scoped to the supplier's contact, with the same fences as vendor_portal_pos. It also explains the two auth paths (grantToken requiring remittance.read, or contactId for operator preview), though it does not explicitly list exclusions or when to prefer a different sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
vendor_portal_revokeA
Widerrufe einen Lieferanten-Portalzugang (US-F03.4): stamps revoked_at so any read of the token now denies as grant_invalid. Fenced to kind=vendor (a non-vendor id returns not_found), so the vendor verb never touches a customer grant. The row is NEVER deleted (the grant history is the revDSG access trail). Revoking an already-revoked grant is a no-op returning the original state.
| Name | Required | Description | Default |
|---|---|---|---|
| actor | No | ||
| grantId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, the description thoroughly discloses critical behaviors: it never deletes the row (preserving audit trail), it is idempotent (revoking an already-revoked grant is a no-op), and it stamps a timestamp that denies access. This is beyond what the schema provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and front-loaded with the main action and effect. It efficiently packs important details about fencing and idempotency into a few sentences. It could be more structured but is not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is fairly complete for a single-action tool: it explains the effect, the fencing, the no-delete policy, and idempotency. However, it lacks explicit parameter semantics and does not mention required fields, but given the tool's simplicity, it is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must explain parameters. It does not mention workspaceId, grantId, actor, or idempotencyKey. The grantId is implicit but not explicitly described. The lack of parameter detail is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action: revoking a vendor portal access (US-F03.4). It uses precise verbs and resources, and explicitly mentions the effect (stamps revoked_at, denies as grant_invalid). This distinguishes it from related tools like portal_grant_revoke and vendor_portal_grants_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool: when revoking a vendor portal access. It also gives an exclusion: 'Fenced to kind=vendor (a non-vendor id returns not_found), so the vendor verb never touches a customer grant.' However, it does not explicitly name alternative tools for revoking other grant types, but the fenced scope implies it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
verify_backupARead-only
Verify a .tillbackup/.tillexport file`s per-table checksums and schema-version compatibility without writing anything, reporting entry count and whether every posted entry balances.
| Name | Required | Description | Default |
|---|---|---|---|
| source | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations include readOnlyHint=true, and the description reinforces that by stating 'without writing anything.' It adds behavioral transparency by listing exactly what the verification covers (per-table checksums, schema-version compatibility, entry count, balance). It doesn't state the actual return payload structure, but no output schema exists, so the description carries reasonable weight.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One long but information-dense sentence. Every clause adds meaning: file types, what is verified, that it is side-effect free, and what it reports. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only verification tool, the description covers purpose, side-effect safety, and the check outputs. It doesn't detail the return schema (e.g., JSON structure of checksum results), but with no output schema provided and the large sibling list, this is a minor omission. The main gap is how to specify which file to verify, though the file-type mention in the description helps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does mention the file type ('.tillbackup/.tillexport'), which gives meaning to the 'source' parameter, but it does not explain what form the source takes (path, uploaded file, URL) or how to reference it. This is minimal but not complete compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Verify') and resource ('.tillbackup/.tillexport' files) and states exactly what it checks: per-table checksums, schema-version compatibility, entry count, and balance. This distinguishes it from sibling backup/restore tools like create_backup, list_backups, delete_backup, and restore_backup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly conveys that this is a validation/verification step for backup/export files, fitting alongside create_backup/list_backups/restore_backup. It doesn't explicitly mention when not to use it or name alternatives like list_restorable_backups, but the context of 'verifying a file' is clear enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voice_buildA
Lerne den Schreibstil aus gesendeten Nachrichten: reads the outbound corpus of one E04 mail account (plus optional E00 documents by id), distils a readable style card, and embeds every exemplar through the LOCAL OP6 runtime, storing locators, vectors and hashes and NEVER an excerpt (Art. 321 index-never-copy). Fewer than 20 sent messages answers corpus_too_small with the honest have/need counts; no installed runtime answers needs_local_runtime, never a cloud fallback; no chosen model answers needs_model_selection. Re-running with a new key rebuilds and supersedes by row, the previous profile retained.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| accountId | Yes | ||
| documentIds | No | ||
| workspaceId | Yes | ||
| idempotencyKey | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden and succeeds: it discloses local-only processing, storage of locators/vectors/hashes, the never-excerpt privacy rule under Art. 321, re-run supersede behavior, and all relevant error conditions. This is unusually transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and information-rich with no filler, front-loading the core purpose. It is structured as one long multi-clause paragraph, which is slightly harder to parse than a list, but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with no annotations, five parameters, and no output schema, the description is remarkably complete: it explains inputs, storage side effects, privacy constraints, failure conditions, idempotency, and what success produces (a readable style card). Missing exact success return shape is minor given the overall coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning for accountId ('one E04 mail account'), documentIds ('optional E00 documents by id'), and idempotencyKey ('re-running with a new key'). workspaceId and name are not elaborated, but the required parameters are at least contextually grounded.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action (learn/build the writing style), a specific resource (outbound corpus of one E04 mail account), and defines what happens: distills a style card and embeds exemplars. It is clearly distinguishable from sibling tools like voice_retrieve or voice_profile_get, which are retrieval-oriented.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit preconditions: at least 20 sent messages, a locally installed OP6 runtime, and a chosen model. It also states explicit failure modes (corpus_too_small, needs_local_runtime, needs_model_selection) and rules out cloud fallback. It does not name alternative tools or say 'use voice_retrieve instead', so it falls just short of the highest bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voice_profile_getBRead-only
Ein Schreibstil-Profil (P5): the readable style card (greeting, sign-off, formality, sentence length, language mix), exemplar count, when it was built, on which model, and stale:true when the sent-mail corpus changed since the build.
| Name | Required | Description | Default |
|---|---|---|---|
| profileId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, aligning with a read operation. The description adds behavioral context by detailing the exact content returned (style card fields, exemplar count, build date, model, staleness flag) and the condition for stale:true, which is valuable beyond the annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single long sentence that mixes German and English, making it somewhat unwieldy. It packs many details but could be restructured for clarity and brevity. The information is dense but not optimally front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description is the sole source of return-value information. It lists many fields but omits types, response structure, error behavior (e.g., profile not found), and how the profileId is obtained. This is incomplete for a get operation without structured output documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the parameters workspaceId and profileId beyond their names. While the names are fairly self-explanatory, the description provides no additional context (e.g., how to obtain profileId or its scope), so it fails to compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (get) and resource (voice profile) and enumerates the returned fields, making the tool's purpose clear. It does not explicitly contrast with sibling tools like voice_profiles_list, but the singular 'profile' and the field list imply it retrieves a specific profile.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this vs. alternatives. The name and description imply usage for retrieving a single profile by ID, but no when-to-use or when-not-to-use conditions are provided, leaving the agent to infer context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voice_profiles_listBRead-only
Alle Schreibstil-Profile (P5), newest first: supersession is by row, so history stays interpretable (a draft references the profile that produced it).
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already marks it as a safe read. The description adds useful behavior: results are newest-first and supersession is by row, so history remains interpretable because drafts reference the profile that produced them. This goes beyond the annotation and clarifies the data model.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence that states the core action (list all profiles, newest first) and then adds a valuable explanatory clause. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with a single obvious parameter and no output schema, the description covers the main behavioral details: resource scope, ordering, and supersession semantics. It leaves the exact return shape unspecified, but that is a minor gap for this simple list operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not mention workspaceId at all. The parameter name is self-explanatory, but the description fails to compensate for the missing schema documentation, leaving the required scope ('profiles in which workspace?') implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description says 'Alle Schreibstil-Profile (P5), newest first', which identifies a specific resource (writing-style profiles, P5) and an ordering. It is clear it lists all profiles, but it does not explicitly name a sibling like voice_profile_get, so it only weakly differentiates from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance. The description explains ordering and supersession semantics but never tells an agent to prefer this over voice_profile_get or voice_build for a given scenario.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
voice_retrieveARead-only
Die ähnlichsten Beispiele zu einem Text (P5, computed at query time): embeds queryText through the local runtime, cosine-ranks the profile exemplars, and returns the top k with bodies read ON DEMAND from their source (E04 mail store, E00 document blob), never from SQLite. A source deleted in its own application is skipped and counted, never fatal; a moved corpus answers stale:true. This is the E06 drafting agent, and every other consumer, reaching the corpus through its only door.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | ||
| profileId | Yes | ||
| queryText | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses significant behavior: bodies are read on demand from source (E04, E00) rather than SQLite, deleted sources are skipped and counted without fatality, and a moved corpus returns stale:true. This provides rich context for the agent's decision-making.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-structured: purpose first, then technical details, then context. It is appropriately sized, though it includes unexplained acronyms (P5, E04, E00) that may confuse. Front-loading is effective.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains behavioral edge cases and the retrieval mechanism, but does not describe the return structure (what fields the top-k items contain) nor fully cover all parameters. With no output schema, this gap leaves the agent unsure of the response format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions queryText (embeds it) and k (returns top k) but fails to explain workspaceId and profileId. It does not provide enough detail to understand all parameters, leaving two required parameters unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: embeds queryText, cosine-ranks profile exemplars, and returns the top k with on-demand bodies. It specifies the resource (profile corpus) and the verb (retrieve similar examples), and distinguishes itself as the sole access door to the corpus, setting it apart from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states that this is the only door to the corpus, implying it should be used for any corpus retrieval. However, it does not explicitly enumerate when not to use it or compare with specific sibling tools like voice_build, leaving some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
void_vendor_billA
Storniere eine Kreditorenrechnung: posts the faithful reversing entry (OR 957a) and flips the bill to void. Never deletes and never edits the original. A DRAFT is simply retired (nothing was booked, so nothing is reversed). A bill with any payment against it is refused with already_settled: reverse the payment first, or 2000 Kreditoren silently carries the difference.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | ||
| reason | No | ||
| workspaceId | Yes | ||
| vendorBillId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses key behaviors: it never deletes or edits the original, it posts a faithful reversing entry, DRAFT bills are retired without reversal, and bills with payments are refused with a specific error code. It also warns about the 2000 Kreditoren account silently carrying the difference, which is critical for an agent to understand the financial impact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and information-dense. It front-loads the primary action and then covers edge cases (DRAFT, settled bills) in a logical order. Every sentence adds value, and the structure is easy to parse for an agent.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the main behavior, edge cases, and error conditions. It lacks an explicit output schema, but the description doesn't need to explain return values if the tool's behavior is clear. It could mention the exact response format or success criteria, but for a void operation, the described behaviors are sufficient for an agent to invoke it correctly. Minor gap: no mention of whether the reversing entry is posted immediately or if there are additional side effects like status changes beyond 'void'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It mentions the core action but does not explain the parameters (date, reason, workspaceId, vendorBillId, idempotencyKey) beyond what the schema names provide. The description implies vendorBillId is the target and idempotencyKey is for idempotency, but it doesn't add detail on date or reason semantics. Baseline 3 is appropriate because the parameter names are self-explanatory and the description gives context for the main one.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: voiding a vendor bill by posting a reversing entry (OR 957a) and flipping the bill to void. It explicitly distinguishes this from deletion or editing, and the sibling list includes related tools like post_vendor_bill and reverse_entry, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use guidance: use for voiding a vendor bill, and it explains the behavior for DRAFT bills (simply retired) and for bills with payments (refused with already_settled, requiring payment reversal first). It also mentions the alternative of reversing the payment first, which helps an agent decide the correct sequence.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
wage_journal_postA
Preview then post the month's externally-computed aggregate wage journal as ONE balanced entry through A02 (source=import). Supply exactly one of lines[] (agent path: each {accountNumber|accountId, debitMinor|creditMinor, costCenter?, description?} in integer Rappen) or fileRef (a stored provider CSV, columns account_number/debit_rappen/credit_rappen/cost_center?/description?, translated through an optional columnMap of canonical->provider header). entryDate is YYYY-MM-DD. Without confirm it returns the full-entry preview and writes NOTHING, without consuming the idempotency key (P8); an unbalanced set is refused at preview time with diffRappen, before A02 is asked. With confirm:true it posts once (requiring the post capability) and records a wage_journal_posts row; a re-post with the same idempotencyKey returns the original posted entry and never posts a second one. A wrong journal is corrected in the payroll system and re-imported, or reversed via A02, never edited. TILL computes no wage: amounts arrive from the provider and are only summed for the balance check. CONSEQUENCE: Posts the month's aggregate wage journal as one balanced entry; a wrong journal is reversed, never edited.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | No | ||
| lines | No | ||
| confirm | No | ||
| fileRef | No | ||
| columnMap | No | ||
| entryDate | Yes | ||
| mappingId | No | ||
| description | No | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so exceptionally. It discloses that without confirm nothing is written and the idempotency key is not consumed, that unbalanced sets are refused at preview with diffRappen, that confirm posts once and records a wage_journal_posts row, and that duplicate posts are idempotent. It also covers capability requirements ('requiring the post capability') and the non-editing correction policy.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but well-organized, with critical information front-loaded (preview then post, balanced entry, source). It uses clear separators for data formats and workflow. Some redundancy exists, notably the 'CONSEQUENCE' section repeating the reversal policy, which could be trimmed without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 10 parameters, nested objects, no output schema, and no annotations, this description is remarkably complete. It explains the full lifecycle (preview, validation, posting, idempotency, correction), return behavior (full-entry preview, diffRappen), and data translation. The only minor gaps are the unexplained parameters noted above, but overall it equips an agent to call the tool correctly in most scenarios.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must supply parameter meaning. It thoroughly explains the two exclusive payloads (lines[] and fileRef), their internal structures, columnMap translation, entryDate format, and confirm/idempotencyKey behavior. However, it does not describe the top-level description, mappingId, or workspaceId parameters, leaving those ambiguous despite the overall strong coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Preview then post the month's externally-computed aggregate wage journal as ONE balanced entry through A02 (source=import).' This clearly differentiates it from sibling posting tools like post_entry and reverse_entry by specifying the wage journal context, the A02 source, and the single balanced entry constraint.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description establishes a clear usage scenario: posting the month's aggregate wage journal from an external provider. It gives explicit workflow guidance (preview first, confirm to post, idempotency) and mentions the correction path ('reversed via A02, never edited'). However, it does not explicitly name alternatives like post_entry or state when NOT to use this tool, so I deduct one point.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
warehouse_archiveA
Soft-archive a warehouse (active=false). Refused with cannot_archive_default while it is the workspace default; refused with location_has_stock or location_in_use when any location under it still holds stock or is referenced by an open stocktake. On success its empty locations are archived too. Deletion is never offered.
| Name | Required | Description | Default |
|---|---|---|---|
| warehouseId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It covers the soft-archive nature, cascading archive of empty locations on success, and refusal conditions. It also notes that deletion is never offered, adding important behavioral context. It doesn't mention idempotency or response behavior, but the core side effects are well described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and structured: it opens with the primary action, then lists refusal conditions, and ends with a note on cascading behavior and deletion. Every sentence adds information without redundancy. It's appropriately sized for the complexity of the operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite no output schema or annotations, the description covers the main operational aspects: the effect (active=false), failure conditions, and side effects on locations. It doesn't describe the return value or error structure, but given the tool's simplicity and the provided refusal reasons, it's reasonably complete for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description should compensate by explaining the parameters. It does not discuss workspaceId, warehouseId, or idempotencyKey beyond the operation context. The parameter names are self-explanatory, but the description adds no specific meaning about their roles or formats, leaving a gap for an agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (soft-archive) and the resource (warehouse), and specifies the resulting state (active=false). It distinguishes itself from sibling tools like warehouse_set_default and warehouse_update by focusing on the archiving behavior and its constraints.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly describes conditions under which the operation is refused (default warehouse, stock in locations, open stocktake references), which informs when the tool is appropriate. It doesn't name specific alternative tools, but the 'Deletion is never offered' note clarifies this is the archive path, not a delete. The usage context is clear, though not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
warehouse_createA
Create a warehouse (a physical or logical site stock lives at). code is 1-20 chars, unique per workspace case-insensitively (duplicate_code); an empty or over-long code is invalid_code. The FIRST warehouse a workspace creates becomes the workspace default; pass isDefault to force it, which demotes the previous default. Plain master data: posts no journal entry.
| Name | Required | Description | Default |
|---|---|---|---|
| city | No | ||
| code | Yes | ||
| name | Yes | ||
| isDefault | No | ||
| postalCode | No | ||
| countryCode | No | ||
| description | No | ||
| workspaceId | Yes | ||
| addressLine1 | No | ||
| addressLine2 | No | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It transparently states code validation rules and error names (duplicate_code, invalid_code), explains the default-demotion side effect, and explicitly notes that this is plain master data with no journal entry. This is strong transparency, missing only details like idempotency semantics or workspace authorization.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences cover the purpose, key validation rules, default behavior, and financial side effect without any filler. Information is front-loaded and each sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create tool, the description covers the most critical behavioral nuances: validation, default assignment, and the lack of posting. It does not mention idempotencyKey usage, but that may be a common pattern across tools. Missing prerequisites like workspace existence are not addressed, but the required workspaceId in the schema partially covers that. Overall it is quite complete for its complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It adds meaning for only two of eleven parameters: code (length, uniqueness, case-insensitivity) and isDefault (default assignment and demotion). All other parameters, including idempotencyKey, workspaceId, and address fields, receive no explanation. This leaves a significant semantic gap for required parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Create a warehouse') and immediately clarifies what a warehouse is ('a physical or logical site stock lives at'). It is clearly distinct from sibling tools like warehouse_get, warehouse_update, and warehouse_archive, which are the obvious alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context about when to use the isDefault parameter and the first-warehouse default behavior. It implicitly steers users toward warehouse_set_default or warehouse_update for changing defaults later, but it does not explicitly name alternatives or exclusion conditions. Since it provides meaningful context without explicit alternatives, it earns a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
warehouse_getARead-only
Read one warehouse by id, archived or not.
| Name | Required | Description | Default |
|---|---|---|---|
| warehouseId | Yes | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, so the description inherits the read-only safety profile. The description adds the fact that archived warehouses are included, which is useful behavioral context beyond the annotation. It does not describe any other behaviors, but with annotations covering the core safety aspect, this is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single short sentence that is front-loaded with the main action and includes the notable detail about archived warehouses. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-by-id tool with two required parameters and no output schema, the description covers the essential purpose and the notable inclusion of archived records. The lack of return value description is acceptable because there is no output schema and the typical expectation is a warehouse object, which is not a major gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description has no parameter details, but the parameter names 'workspaceId' and 'warehouseId' are self-explanatory from the schema itself. The description adds the semantic that it reads one warehouse by ID, which maps directly to the 'warehouseId' parameter. Since the schema provides the names and the description adds minimal extra meaning, a baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Read', the resource 'one warehouse by id', and the scope including archived warehouses. It is specific enough to distinguish from warehouse_list, but does not explicitly name the sibling alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies that this tool is for fetching a single warehouse by ID, and because it includes archived warehouses, it signals when it might be preferred over warehouse_list. However, it does not explicitly state when to use this versus warehouse_list or warehouse_create/update/archive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
warehouse_listARead-only
List the workspace warehouses, default first then by code. Optional active filter (true = only live, false = only archived) and a case-insensitive search over code and name. Accepts a savedViewId (G00 saved-view seam).
| Name | Required | Description | Default |
|---|---|---|---|
| active | No | ||
| search | No | ||
| savedViewId | No | ||
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnlyHint annotation, the description discloses meaningful behavioral traits: default-first ordering, active filter semantics (true vs false), case-insensitive search over code and name, and savedViewId handling. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences deliver the essential information with no filler. Behavior, filters, search, and saved-view support are all front-loaded and clearly organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with one required parameter and no output schema, this description is complete: it covers scope, ordering, filtering, search behavior, and saved-view support. No critical calling information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates well by explaining active, search, and savedViewId semantics. workspaceId is not explicitly described, though its role is strongly implied by 'workspace warehouses' and the required schema field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and explicit resource ('workspace warehouses'), and adds sorting behavior ('default first then by code'). It clearly differentiates this plural listing tool from sibling tools like warehouse_get and warehouse_archive.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear context for use: listing workspace warehouses with an optional live/archived active filter, case-insensitive search, and saved-view support. It does not explicitly name alternatives or state when not to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
warehouse_set_defaultA
Make this warehouse the workspace default, demoting the previous default in the same transaction (exactly one default per workspace). An archived warehouse is refused with warehouse_archived. Omitted-location callers resolve to the default warehouse.
| Name | Required | Description | Default |
|---|---|---|---|
| warehouseId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden itself. It discloses the atomic swap of defaults, the invariant of exactly one default per workspace, the warehouse_archived error, and the resolution behavior for omitted-location callers. It does not cover idempotency semantics or permissions, but the most consequential behaviors are surfaced.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, all informative, with the main effect front-loaded and no filler. The error condition and default-resolution side effect are packed efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a small mutation tool with no output schema, the description covers the core transaction, an edge case, and a behavioral consequence. Missing idempotency-key semantics is the main gap, but the definition is otherwise close to complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds essentially no parameter-level semantics. warehouseId and workspaceId are partially inferable from their names, but idempotencyKey behavior is unexplained, so the description does not compensate for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action: making a warehouse the workspace default, with the transactional nuance of demoting the previous default and keeping exactly one default per workspace. This clearly distinguishes it from siblings like warehouse_archive, warehouse_update, and location_set_default.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended trigger is clear—use this when a warehouse should become the workspace default—and it mentions the archived-warehouse refusal condition. However, it does not explicitly contrast with alternatives or state when not to use it, so some inference is required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
warehouse_updateA
Edit a warehouse through a patch object (name, description, address fields, countryCode). code and the default flag are not changed here (use warehouse_set_default for the default). Only the fields present in patch change.
| Name | Required | Description | Default |
|---|---|---|---|
| patch | No | ||
| warehouseId | Yes | ||
| workspaceId | Yes | ||
| idempotencyKey | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden. It discloses that it is a patch (partial update) operation and that code/default are untouched, which is useful. But it omits behavioral details like what is returned, whether it is idempotent (despite idempotencyKey in schema), or any side effects beyond the field changes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences, front-loaded with the core action and patch object. It efficiently lists editable fields and exclusions without redundancy. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema and moderate complexity, the description explains the patch behavior but lacks details on return value, idempotency semantics, or error handling. It also doesn't mention any validation rules. This is adequate but not comprehensive for a tool with a required idempotency key.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the patch object's fields partially (lists name, description, address fields, countryCode, but not city/postalCode/addressLine1/2 explicitly) and clarifies patching semantics. However, it does not elaborate on workspaceId, warehouseId, or idempotencyKey, which are left to the agent to infer from names and context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it edits a warehouse via a patch object, listing specific fields like name, description, address, countryCode. It explicitly differentiates from the default flag change by pointing to warehouse_set_default, making it distinguishable from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit guidance on when not to use this tool (for changing the default flag) and names the alternative (warehouse_set_default). It also clarifies that only fields present in the patch change, indicating a partial update usage. However, it doesn't explicitly state prerequisites like workspace context beyond the required params.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
whoamiARead-only
Report who this session is in a workspace: the actor, the role it holds, and the exact capabilities that role resolves to. The ONE source a client may use to pre-disable a write control. For the governed agent seat it also carries agentDial, the effective approval-dial level per capability (auto executes, ask drafts), so an agent can predict whether a write will post or wait for approval without holding the owner-only get_agent_dial.
| Name | Required | Description | Default |
|---|---|---|---|
| workspaceId | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint=true annotation already declares the safety profile, so the bar is lower. The description adds genuine behavioral value beyond the annotation: it discloses what the response carries (actor, role, resolved capabilities, and agentDial for the governed seat) and explains how to interpret agentDial for predicting write approval behavior. This does not contradict the annotation — 'report' is consistent with readOnlyHint=true.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in the first sentence, with the unique value proposition following immediately. The text is dense but every clause earns its place — it avoids filler. It is slightly verbose in the agentDial sentence, but the added detail justifies the length. Well-structured overall.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, read-only identity tool with no output schema, the description is thorough. It explains the return contents, the unique use case (pre-disabling write controls), and the relationship to get_agent_dial. No output schema exists, so the description's enumeration of response fields partially compensates. The only omission is the exact response shape, which is minor for this tool type.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description never mentions workspaceId. However, there is only a single required parameter whose name is self-explanatory, so the cost of the omission is low. The description could have added a note (e.g., the workspace context the session is resolved against), but the gap is minor for such an obvious parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource — 'Report who this session is in a workspace: the actor, the role it holds, and the exact capabilities that role resolves to.' This is precise and non-tautological. It also differentiates from siblings by explicitly addressing the relationship to get_agent_dial, noting whoami carries agentDial for the governed seat so the agent need not hold the owner-only get_agent_dial. An agent can immediately distinguish this from the surrounding member/role/dial tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear when-to-use guidance: 'The ONE source a client may use to pre-disable a write control' and the instruction that an agent can 'predict whether a write will post or wait for approval without holding the owner-only get_agent_dial.' This effectively frames the alternative (get_agent_dial) and the condition selecting whoami. It stops short of stating explicit when-not-to-use scenarios, but the context is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
779 tool updates
v0.1.0- First observed
accept_invite - First observed
account_set_tax_default - First observed
accrual_create - First observed
accrual_discard - First observed
accrual_get - First observed
accrual_list - First observed
accrual_post - First observed
accrual_reverse - First observed
advance_move_step - First observed
advance_onboarding_step - First observed
agent_ask - First observed
agent_prose_delete - First observed
agent_trust_summary - First observed
aging_report - First observed
allocate_payment - First observed
apply_qr_match - First observed
approve_drafted_action - First observed
approve_entry - First observed
archive_account - First observed
archive_automation_rule - First observed
archive_bank_account - First observed
archive_contact - First observed
archive_cost_center - First observed
archive_document_template - First observed
archive_field - First observed
archive_item - First observed
archive_role - First observed
archive_workspace - First observed
asset_acquire - First observed
asset_acquisition_summary - First observed
asset_add_capitalisation - First observed
asset_archive - First observed
asset_category_archive - First observed
asset_category_create - First observed
asset_category_get - First observed
asset_category_list - First observed
asset_category_resolve_defaults - First observed
asset_category_update - First observed
asset_create - First observed
asset_depreciation_forecast - First observed
asset_depreciation_method_set_enabled - First observed
asset_depreciation_methods - First observed
asset_depreciation_preview - First observed
asset_depreciation_run_create - First observed
asset_depreciation_run_get - First observed
asset_depreciation_run_list - First observed
asset_depreciation_run_post - First observed
asset_depreciation_run_reverse - First observed
asset_depreciation_schedule - First observed
asset_disposal_get - First observed
asset_disposal_preview - First observed
asset_disposal_summary - First observed
asset_dispose - First observed
asset_end_of_life_list - First observed
asset_get - First observed
asset_ledger_get - First observed
asset_ledger_list - First observed
asset_list - First observed
asset_location_archive - First observed
asset_location_create - First observed
asset_location_get - First observed
asset_location_list - First observed
asset_location_update - First observed
asset_maintenance_log_cancel - First observed
asset_maintenance_log_create - First observed
asset_maintenance_log_get - First observed
asset_maintenance_log_list - First observed
asset_maintenance_log_update - First observed
asset_nbv_summary - First observed
asset_opening_balance - First observed
asset_reconciliation_check - First observed
asset_reconciliation_report - First observed
asset_register_report - First observed
asset_search - First observed
asset_transaction_get - First observed
asset_transaction_list - First observed
asset_transfer - First observed
asset_transfer_history - First observed
asset_update - First observed
attach_receipt - First observed
attention_list - First observed
attention_summary - First observed
balance_sheet - First observed
bank_channel_connect - First observed
bank_channel_directory - First observed
bank_channel_disconnect - First observed
bank_channel_status - First observed
bank_sync - First observed
billing_generate_invoice - First observed
billing_release_time - First observed
billing_unbilled_preview - First observed
billing_wip_report - First observed
bootstrap_workspace - First observed
capture_commit - First observed
capture_discard - First observed
capture_document - First observed
capture_extract - First observed
checklist_abandon - First observed
checklist_get - First observed
checklist_item_complete - First observed
checklist_item_reopen - First observed
checklist_item_skip - First observed
checklist_list - First observed
checklist_start - First observed
checklist_templates - First observed
clear_diagnostics - First observed
close_month - First observed
close_year - First observed
comment_entry - First observed
confirm_field - First observed
confirm_match - First observed
contacts_anonymise - First observed
contacts_import - First observed
contacts_log_activity - First observed
contacts_merge - First observed
contacts_tag - First observed
contacts_timeline - First observed
convert_document - First observed
costing_budget_vs_actual - First observed
costing_drilldown - First observed
costing_pl_list - First observed
costing_project_pl - First observed
create_account - First observed
create_automation_rule - First observed
create_backup - First observed
create_bank_account - First observed
create_contact - First observed
create_cost_center - First observed
create_credit_note - First observed
create_demo_workspace - First observed
create_document - First observed
create_document_template - First observed
create_entry_for_txn - First observed
create_item - First observed
create_payment_batch - First observed
create_recurring_schedule - First observed
create_saved_view - First observed
create_vendor_bill - First observed
create_workspace - First observed
customer_balance - First observed
dashboard_overview - First observed
dashboard_tile - First observed
deals_create - First observed
deals_list - First observed
deals_log_activity - First observed
deals_mark - First observed
deals_move - First observed
deals_to_quote - First observed
deals_update - First observed
define_field - First observed
define_role - First observed
delete_account - First observed
delete_backup - First observed
delete_cost_center - First observed
delete_draft - First observed
delete_item - First observed
delete_saved_view - First observed
delivery_note_create - First observed
delivery_note_issue - First observed
delivery_note_render - First observed
delivery_status - First observed
describe_rate_feed - First observed
detect_anomalies - First observed
disable_automation_rule - First observed
disable_plugin - First observed
discard_demo_workspace - First observed
discard_payment_batch - First observed
discard_testmandant - First observed
dispatch_preview - First observed
dispatch_text_upsert - First observed
draft_generate - First observed
draft_list - First observed
draft_regenerate - First observed
ebill_delivery_status - First observed
ebill_prepare - First observed
ebill_transmit - First observed
egress_self_test - First observed
egress_status - First observed
enable_automation_rule - First observed
enable_plugin - First observed
end_recurring_schedule - First observed
env_copy - First observed
env_create - First observed
env_current - First observed
env_delete - First observed
env_list - First observed
env_reset - First observed
env_status - First observed
env_switch - First observed
expense_claim_approve - First observed
expense_claim_create - First observed
expense_claim_get - First observed
expense_claim_list - First observed
expense_claim_reimburse - First observed
expense_claim_reject - First observed
expense_claim_submit - First observed
expense_line_upsert - First observed
export_journal - First observed
export_statement - First observed
export_statements - First observed
export_vat - First observed
export_workspace - First observed
files_delete - First observed
files_get_content - First observed
files_link - First observed
files_list_linked - First observed
files_new_version - First observed
files_search - First observed
files_set_retention - First observed
files_update - First observed
files_upload - First observed
files_upload_begin - First observed
files_upload_chunk - First observed
files_upload_commit - First observed
flag_entry - First observed
folders_delete - First observed
folders_list - First observed
folders_upsert - First observed
forecast_revenue - First observed
forecast_sales_kpis - First observed
forecast_vs_actual - First observed
forecast_weighted_pipeline - First observed
fx_revaluation - First observed
fx_revaluation_reverse - First observed
general_ledger - First observed
generate_pain001 - First observed
get_agent_dial - First observed
get_agent_session - First observed
get_aging_bucket_config - First observed
get_api_catalog - First observed
get_audit_log - First observed
get_automation_rule - First observed
get_automation_run - First observed
get_bank_account - First observed
get_capture - First observed
get_company_profile - First observed
get_concept - First observed
get_contact - First observed
get_diagnostics - First observed
get_document - First observed
get_document_template - First observed
get_dunning_config - First observed
get_dunning_pdf - First observed
get_dunning_run - First observed
get_ebill_config - First observed
get_entry - First observed
get_exchange_rate - First observed
get_fx_method - First observed
get_item - First observed
get_move_state - First observed
get_onboarding_progress - First observed
get_opening_balances - First observed
get_payment - First observed
get_payment_batch - First observed
get_plugin - First observed
get_plugin_registry_entry - First observed
get_recurring_schedule - First observed
get_sync_contract - First observed
get_vendor_bill - First observed
get_workspace - First observed
gl_archive_account_history - First observed
gl_archive_import - First observed
gl_archive_periods - First observed
gl_archive_preview - First observed
gl_archive_purge - First observed
gl_archive_query - First observed
go_productive - First observed
goods_receipt_accept_lines - First observed
goods_receipt_cancel - First observed
goods_receipt_create - First observed
goods_receipt_get - First observed
goods_receipt_get_config - First observed
goods_receipt_lines_for_match - First observed
goods_receipt_list - First observed
goods_receipt_post - First observed
goods_receipt_preview - First observed
goods_receipt_reject_lines - First observed
goods_receipt_reverse - First observed
goods_receipt_set_config - First observed
goods_receipt_upsert_lines - First observed
hr_absence_cancel - First observed
hr_absence_list - First observed
hr_absence_record - First observed
hr_employee_get - First observed
hr_employee_list - First observed
hr_employee_upsert - First observed
implementation_decision_record - First observed
implementation_parallel_check - First observed
implementation_parallel_declare - First observed
implementation_parallel_status - First observed
implementation_project_close - First observed
implementation_project_create - First observed
implementation_project_get - First observed
implementation_project_list - First observed
implementation_runbook_instantiate - First observed
implementation_signoff_record - First observed
implementation_task_set - First observed
import_camt - First observed
import_exchange_rates - First observed
import_open_items - First observed
import_opening_balances - First observed
income_statement - First observed
install_plugin - First observed
inventory_adjust - First observed
inventory_adjust_analysis - First observed
inventory_adjust_batch - First observed
inventory_adjust_list - First observed
inventory_adjust_reverse - First observed
inventory_alerts - First observed
inventory_anomalies - First observed
inventory_available_serials - First observed
inventory_balance - First observed
inventory_balance_by_location - First observed
inventory_cycle_count_status - First observed
inventory_ensure_default_location - First observed
inventory_get_config - First observed
inventory_lot_trace - First observed
inventory_low_stock - First observed
inventory_move - First observed
inventory_movement_get - First observed
inventory_movement_history - First observed
inventory_movement_list - First observed
inventory_on_hand_by_lot - First observed
inventory_reason_archive - First observed
inventory_reason_create - First observed
inventory_reason_get - First observed
inventory_reason_list - First observed
inventory_reason_update - First observed
inventory_reconciliation_check - First observed
inventory_reconciliation_report - First observed
inventory_reorder_candidates - First observed
inventory_set_config - First observed
inventory_slow_movers - First observed
inventory_stock_position - First observed
inventory_stocktake_approve_lines - First observed
inventory_stocktake_cancel - First observed
inventory_stocktake_commit - First observed
inventory_stocktake_count - First observed
inventory_stocktake_create - First observed
inventory_stocktake_get - First observed
inventory_stocktake_list - First observed
inventory_stocktake_report - First observed
inventory_stocktake_request_recount - First observed
inventory_transfer - First observed
inventory_valuation_create - First observed
inventory_valuation_get - First observed
inventory_valuation_layers - First observed
inventory_valuation_list - First observed
inventory_valuation_method_history - First observed
inventory_valuation_method_set_enabled - First observed
inventory_valuation_methods - First observed
inventory_valuation_opening - First observed
inventory_valuation_post - First observed
inventory_valuation_preview - First observed
inventory_valuation_report - First observed
inventory_valuation_reverse - First observed
inventory_valuation_set_default - First observed
inventory_valuation_set_item_method - First observed
inventory_valuation_status - First observed
invite_member - First observed
issue_credit_note - First observed
issue_dunning_run - First observed
issue_invoice - First observed
item_categories_delete - First observed
item_categories_list - First observed
item_categories_upsert - First observed
item_set_tracking_mode - First observed
landed_cost_allocate_confirm - First observed
landed_cost_allocate_preview - First observed
landed_cost_get - First observed
landed_cost_list - First observed
landed_cost_reverse - First observed
landed_cost_voucher_create - First observed
ledger_qa - First observed
list_accounts - First observed
list_agent_sessions - First observed
list_automation_rules - First observed
list_automation_runs - First observed
list_backups - First observed
list_bank_accounts - First observed
list_bank_statements - First observed
list_captures - First observed
list_concepts - First observed
list_contacts - First observed
list_cost_centers - First observed
list_dispatches - First observed
list_document_templates - First observed
list_documents - First observed
list_drafted_actions - First observed
list_dunning_runs - First observed
list_exchange_rates - First observed
list_feedback - First observed
list_field_defs - First observed
list_field_values - First observed
list_items - First observed
list_journal - First observed
list_members - First observed
list_open_items - First observed
list_payable - First observed
list_payment_batches - First observed
list_payments - First observed
list_payroll_handoffs - First observed
list_period_locks - First observed
list_plugins - First observed
list_reconciliation - First observed
list_recurring_schedules - First observed
list_restorable_backups - First observed
list_roles - First observed
list_saved_views - First observed
list_unmatched_incoming - First observed
list_vendor_bills - First observed
list_workspaces - First observed
location_archive - First observed
location_create - First observed
location_get - First observed
location_list - First observed
location_set_default - First observed
location_tree - First observed
location_update - First observed
lock_period - First observed
lot_archive - First observed
lot_create - First observed
lot_get - First observed
lot_list - First observed
lot_search - First observed
lot_set_status - First observed
lot_update - First observed
mail_accounts_list - First observed
mail_connect - First observed
mail_draft_write - First observed
mail_drafts_list - First observed
mail_reindex - First observed
mail_thread_get - First observed
mail_threads_list - First observed
mark_batch_paid - First observed
match_bill - First observed
match_qr_payment - First observed
match_status_for_bill - First observed
match_three_way_create - First observed
match_three_way_evaluate - First observed
match_three_way_exceptions - First observed
match_three_way_get - First observed
match_three_way_list - First observed
match_three_way_override - First observed
match_three_way_reverse - First observed
migration_abandon_plan - First observed
migration_apply_map_template - First observed
migration_check_step - First observed
migration_close_plan - First observed
migration_commit_step - First observed
migration_create_plan - First observed
migration_create_testmandant - First observed
migration_declare_control_total - First observed
migration_diff_testmandant_to_live - First observed
migration_discover_source - First observed
migration_export_check - First observed
migration_get_check - First observed
migration_get_extraction_guide - First observed
migration_get_manifest - First observed
migration_get_map - First observed
migration_get_plan - First observed
migration_get_testmandant - First observed
migration_list_checks - First observed
migration_list_extraction_guides - First observed
migration_list_locale_packs - First observed
migration_list_map_templates - First observed
migration_list_plans - First observed
migration_list_source_adapters - First observed
migration_preview_step - First observed
migration_readiness - First observed
migration_record_approval - First observed
migration_rollback_step - First observed
migration_save_map_template - First observed
migration_set_manifest - First observed
migration_set_manifest_item - First observed
migration_set_map - First observed
migration_set_scope - First observed
migration_suggest_map - First observed
migration_trial_load_step - First observed
migration_waive_control - First observed
month_end_checklist - First observed
notifications_archive - First observed
notifications_deliver - First observed
notifications_list - First observed
notifications_list_preferences - First observed
notifications_mark_all_read - First observed
notifications_mark_read - First observed
notifications_run_digest - First observed
notifications_set_preference - First observed
onboard_client - First observed
override_qr_match - First observed
pause_recurring_schedule - First observed
payment_batch_transmit - First observed
payroll_handoff_export - First observed
pipeline_stages_upsert - First observed
pipelines_upsert - First observed
po_amendment_apply - First observed
po_amendment_cancel - First observed
po_amendment_preview - First observed
po_amendment_reject - First observed
po_amendment_start - First observed
po_amendment_submit - First observed
po_amendment_update_lines - First observed
po_cancel - First observed
po_close_short - First observed
po_get - First observed
po_list - First observed
po_open_lines - First observed
po_revise - First observed
po_send - First observed
po_upsert - First observed
po_version_diff - First observed
po_version_get - First observed
po_version_list - First observed
portal_grant_create - First observed
portal_grant_list - First observed
portal_grant_revoke - First observed
portal_grant_send - First observed
portal_quote_accept - First observed
portal_resolve - First observed
post_entry - First observed
post_fx_revaluation - First observed
post_vendor_bill - First observed
prepare_feedback - First observed
prepare_period - First observed
preview_bank_opening_balance - First observed
preview_document_template - First observed
preview_feedback - First observed
preview_open_items - First observed
preview_opening_import - First observed
preview_payment - First observed
preview_plugin_install - First observed
price_lists_delete - First observed
price_lists_get - First observed
price_lists_list - First observed
price_lists_set_price - First observed
price_lists_unset_price - First observed
price_lists_upsert - First observed
price_resolve - First observed
procurement_anomalies - First observed
procurement_grir_clearing - First observed
procurement_landed_cost_variance - First observed
procurement_match_status - First observed
procurement_open_commitments - First observed
procurement_po_cycle - First observed
procurement_po_history - First observed
procurement_requisition_pipeline - First observed
procurement_spend_summary - First observed
procurement_supplier_scorecard - First observed
project_budget_actual - First observed
project_create - First observed
project_delete - First observed
project_get - First observed
project_list - First observed
project_phase_add - First observed
project_phase_done - First observed
project_phase_update - First observed
project_set_status - First observed
project_update - First observed
propose_dunning_run - First observed
provision_create - First observed
provision_discard - First observed
provision_get - First observed
provision_list - First observed
provision_post - First observed
provision_release - First observed
provision_release_reverse - First observed
provision_reverse - First observed
quotes_accept - First observed
quotes_convert - First observed
quotes_create - First observed
quotes_decline - First observed
quotes_expire_sweep - First observed
quotes_get - First observed
quotes_list - First observed
quotes_revise - First observed
quotes_send - First observed
quotes_update - First observed
rate_card_end - First observed
rate_card_list - First observed
rate_card_upsert - First observed
receipt_record - First observed
record_exchange_rate - First observed
record_expense - First observed
record_incoming_credit - First observed
record_payment - First observed
refresh_plugin_compat - First observed
reject_drafted_action - First observed
reopen_month - First observed
reports_delete - First observed
reports_duplicate - First observed
reports_list - First observed
reports_preview - First observed
reports_run - First observed
reports_runs - First observed
reports_save - First observed
reports_schedule - First observed
reports_sources - First observed
reports_update - First observed
requisition_approve - First observed
requisition_cancel - First observed
requisition_close - First observed
requisition_convert_to_po - First observed
requisition_get - First observed
requisition_list - First observed
requisition_my_pending_approvals - First observed
requisition_reject - First observed
requisition_return - First observed
requisition_submit - First observed
requisition_upsert - First observed
restore_backup - First observed
resume_recurring_schedule - First observed
retainer_burndown - First observed
retainer_close - First observed
retainer_create - First observed
retainer_generate_invoice - First observed
retainer_list - First observed
retainer_run_due - First observed
retainer_update - First observed
retry_automation_run - First observed
reverse_entry - First observed
reverse_payment - First observed
review_bank_txn - First observed
review_status - First observed
revoke_member - First observed
run_due_automations - First observed
run_due_recurring - First observed
runtime_catalog - First observed
runtime_select - First observed
runtime_status - First observed
sales_order_backorders - First observed
sales_order_cancel - First observed
sales_order_confirm - First observed
sales_order_create - First observed
sales_order_from_quote - First observed
sales_order_get - First observed
sales_order_invoice - First observed
sales_order_list - First observed
save_draft - First observed
search_global - First observed
search_plugin_registry - First observed
send_dunning_run - First observed
send_invoice - First observed
serial_archive - First observed
serial_create - First observed
serial_create_bulk - First observed
serial_get - First observed
serial_list - First observed
serial_search - First observed
serial_set_status - First observed
serial_update - First observed
set_agent_dial - First observed
set_aging_bucket_config - First observed
set_bank_opening_balance - First observed
set_bank_sync_schedule - First observed
set_camt_matching - First observed
set_creditor_bank_profile - First observed
set_creditor_profile - First observed
set_default_document_template - First observed
set_diagnostics - First observed
set_dunning_config - First observed
set_ebill_config - First observed
set_field_value - First observed
set_fiscal_config - First observed
set_fx_method - First observed
set_opening_balances - First observed
set_qr_auto_apply - First observed
set_role - First observed
set_vat_method - First observed
set_write_off_threshold - First observed
sign_requests_complete - First observed
sign_requests_create - First observed
sign_requests_delete_draft - First observed
sign_requests_get - First observed
sign_requests_list - First observed
sign_requests_record_event - First observed
sign_requests_send - First observed
sign_requests_withdraw - First observed
stock_location_upsert - First observed
stock_low_stock - First observed
stock_move - First observed
stock_on_hand - First observed
stock_run_valuation - First observed
stock_stocktake_commit - First observed
stock_stocktake_count - First observed
stock_stocktake_open - First observed
stock_stocktake_report - First observed
stock_valuation_report - First observed
suggest_matches - First observed
suggest_payment_matches - First observed
supplier_performance_alerts - First observed
supplier_performance_explain - First observed
supplier_performance_rank - First observed
supplier_performance_trend - First observed
supplier_price_list - First observed
supplier_price_upsert - First observed
supplier_scorecard_get - First observed
sync_artifact_read - First observed
sync_publish_disable - First observed
sync_publish_enable - First observed
sync_stream_read - First observed
sync_stream_status - First observed
tasks_cancel - First observed
tasks_complete - First observed
tasks_create - First observed
tasks_list - First observed
tasks_reminders_due - First observed
tasks_snooze - First observed
tasks_update - First observed
tax_provision_preview - First observed
time_approve - First observed
time_delete - First observed
time_list - First observed
time_lock - First observed
time_log - First observed
time_resolve_rate - First observed
time_start - First observed
time_stop - First observed
time_submit - First observed
time_update - First observed
transition_document - First observed
trial_balance - First observed
unarchive_account - First observed
unarchive_bank_account - First observed
unarchive_contact - First observed
unarchive_cost_center - First observed
unarchive_item - First observed
uninstall_plugin - First observed
unlock_period - First observed
update_account - First observed
update_automation_rule - First observed
update_bank_account - First observed
update_company_profile - First observed
update_contact - First observed
update_document - First observed
update_document_template - First observed
update_item - First observed
update_recurring_schedule - First observed
update_saved_view - First observed
vat_annual_reconciliation - First observed
vat_code_deactivate - First observed
vat_code_reactivate - First observed
vat_code_upsert - First observed
vat_codes - First observed
vat_config - First observed
vat_configure - First observed
vat_export_ech0217 - First observed
vat_mark_filed - First observed
vat_periods - First observed
vat_preview - First observed
vat_return - First observed
vat_saldo_declaration_basis - First observed
vat_saldo_eligibility - First observed
vat_saldo_generations - First observed
vat_seed_defaults - First observed
vat_settlement_list - First observed
vat_settlement_post - First observed
vat_settlement_preview - First observed
vat_settlement_reverse - First observed
vendor_portal_grant - First observed
vendor_portal_grants_list - First observed
vendor_portal_pos - First observed
vendor_portal_remittance_create - First observed
vendor_portal_remittances - First observed
vendor_portal_revoke - First observed
verify_backup - First observed
voice_build - First observed
voice_profile_get - First observed
voice_profiles_list - First observed
voice_retrieve - First observed
void_vendor_bill - First observed
wage_journal_post - First observed
warehouse_archive - First observed
warehouse_create - First observed
warehouse_get - First observed
warehouse_list - First observed
warehouse_set_default - First observed
warehouse_update - First observed
whoami
TDQS
Scored across 779 tools
Hundreds of tools repeat the same concept under different families: stock_move/inventory_move, stock_on_hand/inventory_balance/inventory_stock_position, stock_stocktake_open/inventory_stocktake_create, stock_run_valuation/inventory_valuation_create. Even with detailed descriptions, an agent cannot reliably distinguish legacy and current surfaces or overlapping read models.
All names are snake_case, but the pattern is inconsistent: list_accounts and list_contacts sit beside tasks_list and reports_runs; get_contact and get_entry sit beside asset_get and warehouse_get. Domain prefixes, verb-first names, noun-first names, and German/English mixes are used interchangeably.
779 tools is an extreme count for any MCP server. Even a full Swiss ERP surface would be better exposed as focused module servers or a much smaller curated tool set; this volume is not navigable.
The tool surface is extraordinarily comprehensive: ledgers, VAT, receivables, payments, banking, inventory, fixed assets, procurement, projects, time, payroll, migrations, plugins, automation, and reporting all have full lifecycle coverage. I cannot identify an obvious domain gap, though the coverage is achieved through massive redundancy.
Maintenance
Related MCP Connectors
Headless API-first double-entry accounting & bookkeeping engine. 84 MCP tools over HTTP.
Hosted MCP server for Mini Accountant: invoices, expenses, customers, analytics, tax estimates.
Open-source AI accounting skills verified by licensed accountants (tax, VAT, payroll).
- Frihet ERPOAuthio.frihet
AI-native ERP MCP: ES/EU fiscal compliance (VeriFactu/TicketBAI/Facturae), invoicing, tax, banking
Related MCP Servers
- FlicenseBqualityDmaintenanceMCP server for DACH accounting automation. Connect AI assistants to sevDesk and Lexoffice — create invoices, manage contacts, handle bookings and vouchers for German-speaking businesses.1556 npm-
- AlicenseNot gradedqualityAmaintenanceA Model Context Protocol (MCP) server that keeps the books for your personal and business finances using double-entry accounting — driven entirely from an LLM.220 npmMIT
- AlicenseAqualityCmaintenanceMCP server for Spanish accounting for freelancers and SMEs, enabling AI agents to issue invoices, OCR expense PDFs, reconcile bank transactions, and prepare quarterly VAT (Modelo 303).23MIT
- AlicenseNot gradedqualityAmaintenanceMCP server for agent-first double-entry bookkeeping for Dutch SMEs, supporting VAT, Peppol BIS 3.0 e-invoicing, and local-first SQLite storage with full audit logging and deterministic JSON output.1Apache 2.0