finance
Server Details
Personal finance for AI agents: accounts, budgets, goals, 9-strategy debt payoff, reports. OAuth 2.1
- Status
- Healthy
- Uptime
- 28.6% over 47 days
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-03-26
- URL
TDQS
Scored across 179 tools
179 tools span many domains, and within analytics/accounts/debt there are multiple overlapping 'get_*' summaries (cash flow, income statement, spending, period summary, daily averages, financial ratios, balance sheet, net worth) whose boundaries are subtle. Descriptions help, but the volume and similar names create real misselection risk.
Nearly all tools use snake_case verb_noun (create_account, list_transactions, update_budget, delete_goal) with consistent prefixes like get_/list_/create_/update_/delete_. Minor variations such as bulk_ or confirm_ still fit the same predictable pattern.
179 tools is an extreme mismatch for any single MCP server; the rubric caps 50+ tools at score 1. The set spans dozens of domains, far beyond a coherent, well-scoped surface.
The server covers account, transaction, budget, goal, category, tag, family, forum, Plaid, subscription, scenario, support, monitoring, and gamification lifecycles with CRUD, analytics, import/export, and edge-case operations. No obvious dead ends for a personal-finance assistant.
Available Tools
179 toolsaccept_family_invitationFamily: Accept family invitationAInspect
Accept a pending family invitation.
Args:
invitation_id: ID of the invitation to accept
Returns:
Confirmation with group details.
| Name | Required | Description | Default |
|---|---|---|---|
| invitation_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already signal that this is not read-only or idempotent, so the description is not burdened with the full safety profile. It does add the return behavior ('Confirmation with group details'), but does not disclose side effects such as the invitation being consumed or the user joining the family group.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and uses a compact Args/Returns structure. Every sentence contributes necessary information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter action the description gives the parameter meaning and return type, but it omits practical context such as how to obtain the invitation_id and what state transition occurs upon acceptance. Given this is a mutating action with no output schema, a bit more behavioral context would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the description compensates by explaining that invitation_id is 'ID of the invitation to accept.' This adds semantic meaning beyond the schema's bare integer field and is sufficient for callers.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Accept') and resource ('pending family invitation'), clearly distinguishing this from related sibling tools like decline_family_invitation or revoke_family_invitation. Even without reading the schema, an agent knows exactly what action this performs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for pending invitations, but it does not explicitly state when to choose acceptance over declining or revoking, nor does it mention prerequisites such as the invitation belonging to the current user. Usage context is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
accept_plaid_balanceAccounts: Accept Plaid-reported balanceAInspect
Accept Plaid's current balance as correct, clearing any mismatch warning.
This doesn't change the balance value (unlike a balance override)
— it just acknowledges the current balance is accurate and removes the
warning flag on the account. Plaid-linked accounts only.
Args:
account_id: The Plaid-linked account
Returns:
Updated account snapshot
| Name | Required | Description | Default |
|---|---|---|---|
| account_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only give generic write/non-destructive flags; the description adds real substance by disclosing that it does NOT change the balance value and only removes a warning flag, so the name doesn't mislead. It doesn't address idempotency or failure modes, which is the remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then supporting context, then Args/Returns. The docstring-style Args/Returns blocks add slight overhead, but the Returns note is justified given no output schema exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 1-param tool with no output schema, the description covers scope, behavioral effect, and the return value ('Updated account snapshot'). Minor omissions like behavior on non-Plaid accounts and idempotency keep it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load — and it adds meaningful constraint by defining account_id as 'The Plaid-linked account', not just any account. It lacks format/type detail (integer ID), but compensates for the coverage gap reasonably well.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource — accepting Plaid's reported balance as correct and clearing the mismatch warning — and explicitly distinguishes itself from the sibling override_account_balance ('unlike a balance override'). An agent can tell exactly what this does and how it differs from the value-changing alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly scopes applicability ('Plaid-linked accounts only') and contrasts with the balance-override alternative, implying when to use each. It stops short of an explicit when-not-to-use statement or error handling for non-Plaid accounts, but the routing intent is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
activate_themeGamification: Activate themeAInspect
Activate a color theme.
Args:
theme_id: Theme identifier (light, dark, emerald_banking, frost_glass, carbon)
Returns:
Confirmation with the activated theme
| Name | Required | Description | Default |
|---|---|---|---|
| theme_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is a non-read-only, non-destructive mutation. The description adds that it returns a confirmation with the activated theme, which is useful, but it does not disclose whether the change persists, affects only the current session, or validates the theme_id beyond the listed values.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured with an action verb up front and clearly labeled Args and Returns sections. Every sentence contributes needed information without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no output schema, the description covers the action, parameter values, and return value. It could mention whether activation affects the current user or the whole workspace, but given the simple scope, the missing detail is acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides no description for theme_id, but the description compensates by listing the valid identifiers (light, dark, emerald_banking, frost_glass, carbon). This gives the agent concrete options that the schema lacks, although it stops short of explaining the visual differences or return format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Activate a color theme', naming the exact verb and resource. It also enumerates the specific theme IDs in the Args section, making the tool's scope unambiguous and distinct from the sibling list_themes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used to change the current theme, but it provides no explicit guidance on when to use it versus alternatives, nor any prerequisites like listing available themes first. Usage is self-evident but not explicitly contextualized.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
apply_statementStatements: Apply parsed statement fields to an accountAInspect
Apply user-confirmed statement fields to one of the user's accounts.
Premium-only. Only the fields passed are written; each is re-validated
server-side (same bounds as the web + REST paths) under a row lock.
Args:
account_id: The account to update (the caller's own, active).
balance: New balance.
interest_rate: APR percent, 0-100, up to 3 decimals.
credit_limit: Credit limit (>= 0).
minimum_payment: Minimum payment (>= 0).
institution: Institution name (max 200 chars).
original_balance: Original loan/balance amount (>= 0).
term_months: Loan term in months (1-600).
loan_start_date: Loan start date, YYYY-MM-DD.
loan_end_date: Loan maturity/payoff date, YYYY-MM-DD.
Returns:
``{"success": True, "account_id": ..., "updated_fields": [...],
"new_balance": ...}`` on success, or ``{"error": "..."}`` (not
Premium, account not found, bad value, or nothing to apply).
| Name | Required | Description | Default |
|---|---|---|---|
| balance | No | ||
| account_id | Yes | ||
| institution | No | ||
| term_months | No | ||
| credit_limit | No | ||
| interest_rate | No | ||
| loan_end_date | No | ||
| loan_start_date | No | ||
| minimum_payment | No | ||
| original_balance | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare it is a non-readonly, non-idempotent, non-destructive mutation, but the description adds real depth: only passed fields are written (partial update), values are re-validated server-side under a row lock (concurrency/validation), and it enumerates error modes (not Premium, account not found, bad value, nothing to apply). It stops short of discussing reversibility or auditing, so a 4 rather than 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-organized with front-loaded purpose followed by an Args block and return shape; the Args lines are justified because the schema carries no descriptions. Slightly list-like but no wasted sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter mutation with no output schema, it covers the essentials: gating (Premium), partial-update semantics, validation, all parameter constraints, and both success and error return shapes. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries full burden and delivers: it documents all ten fields with bounds and formats (interest_rate 0-100 up to 3 decimals, credit_limit >= 0, institution max 200 chars, term_months 1-600, dates as YYYY-MM-DD), which is meaningfully more than the bare type/default schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('apply') + resource ('user-confirmed statement fields') + scope ('to one of the user's accounts'). The qualifier 'user-confirmed' distinguishes it from the sibling parse_statement, which only parses, so an agent can tell the two apart without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'user-confirmed statement fields' implies this runs after parse_statement and after user confirmation, and 'Premium-only' is a clear gating condition. However, it never names the sibling tool to call first nor states when-not to use it (e.g., versus update_account or override_account_balance), leaving the workflow inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bulk_categorize_transactionsTransactions: Bulk categorize transactionsAInspect
Assign a category to multiple transactions at once.
Args:
transaction_ids: List of transaction IDs to categorize
category_name: Category name to assign (case-insensitive)
Returns:
Count of updated transactions.
| Name | Required | Description | Default |
|---|---|---|---|
| category_name | Yes | ||
| transaction_ids | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is a mutating call, and the description adds useful behavioral context with 'Returns: Count of updated transactions' and the case-insensitive nature of category_name. It does not disclose behavior for invalid category names, nonexistent transaction IDs, or whether existing categories are overwritten.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, uses a clear Args/Returns structure, and contains no filler or meaningless repetition. Every sentence contributes useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with only two required parameters and no output schema, the description covers purpose, parameters, and return value adequately. It omits edge-case behavior such as partial failures or category existence requirements, but a caller has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the parameters, and it does: transaction_ids as the list to categorize and category_name as the case-insensitive category to assign. Both required parameters receive meaningful semantic explanations, though more detail about validation or constraints would be helpful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action: 'Assign a category to multiple transactions at once.' This identifies the verb, resource, and batch scope. However, it does not explicitly differentiate the tool from similar siblings such as bulk_delete_transactions or recategorize_similar_transactions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'at once' implies this tool is for batch categorization, which gives some usage context. Yet the description provides no explicit guidance about when to prefer it over single-transaction update tools or alternatives like recategorize_similar_transactions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bulk_delete_transactionsTransactions: Bulk delete transactionsADestructiveInspect
Delete multiple transactions at once with per-row rule-21 semantics.
Each row gets the same treatment as a single-transaction delete:
manual transactions are soft-deleted (balance reversed, restorable
for 7 days); Plaid-originated transactions
are hard-deleted (restoring them after Plaid has moved on creates
ledger drift). Future-dated (unapplied) rows skip balance reversal.
IDs that are missing or belong to another user are skipped and
reported; those rows are left untouched.
Args:
transaction_ids: List of transaction IDs to delete (max 100)
Returns:
Per-id results with `undoable` flags, counts, and a note about
the 7-day undo window for soft-deleted rows.
| Name | Required | Description | Default |
|---|---|---|---|
| transaction_ids | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes far beyond the annotations by disclosing the core behaviors: soft-delete with balance reversal and 7-day restorability for manual rows, hard-delete for Plaid rows and the drift rationale, skipped balance reversal for future-dated rows, and silent skipping/reporting of missing or foreign-user IDs. These are exactly the mutation consequences an agent needs and none contradict the destructive/non-idempotent hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then uses clear Args/Returns structure. The prose is slightly long but each sentence conveys a distinct rule; little is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return contract (per-id results, undoable flags, counts, 7-day undo note) and fully covers the destructive semantics, authorization edge cases, and edge-case row handling. Nothing essential to correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the single array parameter carries no schema description, but the description supplies the crucial 'max 100' bound the schema omits and clarifies the parameter is a list of transaction IDs. That is real added meaning over an otherwise bare integer array.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (delete multiple transactions) plus the operational scope ('at once'), which cleanly separates it from the single-row delete_transaction sibling. An agent can identify the tool's function without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The bulk/at-once framing implies when this tool applies, and the per-row semantics inform decisions, but it never names the alternative (delete_transaction for single rows, undo_transaction for recovery) or states when-not to use it. Usage is inferred rather than instructed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
cancel_subscriptionSubscription: Cancel subscription at period endAInspect
Cancel the user's paid subscription. Premium access continues
until the end of the current billing period.
Returns ``{"success": True, ...}`` or ``{"error": ...}``.
No-op for users on Free, family-shared, or trialing-without-card.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds useful behavioral detail beyond the sparse annotations: cancellation takes effect at period end, the tool returns a success/error envelope, and it is a no-op for certain subscription states. This meaningfully helps an agent predict outcomes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the primary action, followed by the most important behavioral caveat, return format, and no-op conditions. Every sentence adds necessary information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with no output schema, the description is complete: it states the effect, the duration of continued access, the return shape, and the edge cases where the call does nothing. An agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so there is nothing for the schema to cover or the description to explain. Per the baseline for zero-parameter tools, this is a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: cancel the user's paid subscription. It also clarifies the key timing detail (premium access continues until the end of the current billing period), which distinguishes it from an immediate cancellation and from sibling tools like reactivate_subscription.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when the tool is appropriate and explicitly lists no-op cases (Free, family-shared, trialing-without-card). It does not explicitly name alternative tools such as reactivate_subscription or open_billing_portal, which keeps it from a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_payment_coverageAnalytics: Check payment coverageARead-onlyIdempotentInspect
Will upcoming scheduled payments clear? Per-account coverage projection.
For every bill-paying account: the known scheduled payments in the
window (recurring templates — including a brand-new template's first
occurrence — plus Plaid debt due dates), scheduled deposits
(paychecks), a date-ordered projected-balance timeline, the projected
minimum balance, and a status per account:
- ``ok`` — everything clears with cushion to spare
- ``below_threshold`` — clears, but dips under the account's
low-balance cushion (``cushion_shortfall`` says by how much)
- ``overdraft_risk`` — projected NEGATIVE (``overdraft_shortfall``
is the amount needed to cover)
Answers questions like "will I have enough when X hits?", and shows
whether a new recurring withdrawal or scheduled payment would clear.
Shortfall entries carry the amount and the date of the low point.
Covering a shortfall (e.g. a transfer from savings) happens at the
user's bank; Zoninga moves no money.
``unattributed_bills`` are Plaid debt-account due dates that can't be
tied to a specific funding account; they are measured against
``checking_total_balance`` rather than any single account.
Args:
days: look-ahead window in days (default 14, max 90)
Returns:
Per-account coverage entries + overall_status + a summary line.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (readOnly, idempotent, non-destructive), and the description adds substantial behavior on top: the exact status vocabulary (ok / below_threshold / overdraft_risk with the meaning of cushion_shortfall and overdraft_shortfall), that shortfall entries carry amount and low-point date, and the important scope limit that 'Zoninga moves no money' so covering a shortfall happens at the user's bank. It also explains the edge case of unattributed_bills being measured against checking_total_balance rather than a single account.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The question-first opening and the terse status bullets are well front-loaded and every status definition earns its place. The explicit 'Args:' / 'Returns:' block is partly redundant with the schema and the section on unattributed_bills is wordy, so it is dense but slightly longer than necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by stating the return shape (per-account coverage entries + overall_status + summary line) and by defining the individual fields an agent must interpret. Combined with the documented statuses and window parameter, an agent has everything needed to call and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter is bare in the schema (integer with default 14, no description), so the description must carry the burden. It does: 'days: look-ahead window in days (default 14, max 90)' supplies both the meaning and the 90-day upper bound the schema omits.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource — project per-account coverage of upcoming scheduled payments, bills, deposits and Plaid debt due dates — and even supplies the natural-language question it answers ('will I have enough when X hits?'). It is highly distinctive in what it computes, but it never explicitly contrasts itself with the many adjacent analytics siblings such as forecast_cash_flow, get_account_health or get_cash_flow, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is clear: it answers 'will I have enough when X hits?' and can show whether a new recurring withdrawal or scheduled payment would clear. That is a concrete when-to-use signal, but no when-NOT-to-use or named alternative is given, so the agent gets no explicit routing guidance against forecast_cash_flow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
commit_scenarioScenarios: Commit a scenario to real accountsAInspect
Commit a scenario — create its add items as real accounts.
Turns the scenario's add-debt / add-asset changes into real Account
rows and archives any sell/remove targets. Use ONLY after confirming
with the user that they've actually gone through with the change and
the numbers are still correct. Enforces the account limit.
Args:
scenario_id: The scenario to commit.
Returns:
{"success": True, "created_accounts": [{id, name}], "scenario": {...}}
or {"error": "..."}.
| Name | Required | Description | Default |
|---|---|---|---|
| scenario_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes beyond the annotations by disclosing concrete side effects: it creates real Account rows, archives sell/remove targets, and enforces the account limit. This is particularly valuable given that the annotations only provide negative hints (readOnlyHint=false, destructiveHint=false, etc.) and no safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: a one-line summary followed by a detailed behavior sentence, an explicit precondition, and clearly separated Args/Returns sections. Every sentence earns its place with no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description appropriately covers the return shape. It also explains the core side effects, the user-confirmation prerequisite, and the account-limit constraint. This is sufficient for an agent to decide when and how to invoke the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description needs to compensate, but the Args section only restates the parameter as 'The scenario to commit,' which is nearly tautological. It does not explain how to obtain or validate a scenario_id, nor what constitutes a valid scenario for this operation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource ('Commit a scenario') and explains the concrete outcome: creating add-debt/add-asset items as real Account rows and archiving sell/remove targets. This distinguishes it from sibling tools like create_scenario, update_scenario, or delete_scenario without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use this tool ONLY after confirming with the user that the change actually happened and the numbers are still correct. This is a clear conditional usage guideline, though it does not name alternative tools or explicitly say when not to use the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
confirm_debt_payment_matchDebt Payments: Confirm debt-payment matchAInspect
Confirm a pending debt payment match — creates the corresponding
debt-account transaction, increments the rule's confirmed_count, and
auto-enables the rule once the threshold (default 2) is reached.
Args:
debt_payment_id: ID of the pending DebtPayment record to confirm
Returns:
Updated debt payment details
| Name | Required | Description | Default |
|---|---|---|---|
| debt_payment_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=false), the description discloses the exact side effects: creating a transaction, incrementing confirmed_count, and auto-enabling the rule after a threshold (default 2). This adds meaningful behavioral context that annotations alone do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, with the key behavioral summary front-loaded in the first sentence, followed by clearly labeled Args and Returns sections. Every sentence earns its place without redundant fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter mutation tool with no output schema, the description provides the required input semantics, the side effects to expect, and the return value ('Updated debt payment details'). The threshold behavior is also disclosed, so an agent has enough context to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only the parameter name and type with no description, so the description must compensate. The 'Args' section explains that debt_payment_id is 'ID of the pending DebtPayment record to confirm', adding the necessary semantic meaning for correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Confirm') and resource ('pending debt payment match'), and spells out the main effect: creating the corresponding debt-account transaction. This clearly distinguishes it from sibling tools like reject_debt_payment_match and list_debt_payments.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states it is for confirming a pending debt payment match, giving the agent a clear triggering context. It does not explicitly mention alternatives or when not to use the tool, but the 'pending' qualifier provides enough guidance relative to sibling confirm/reject tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contribute_to_goalGoals: Contribute to goalAInspect
Add a contribution to a financial goal.
Args:
goal_id: The goal ID to contribute to
amount: Contribution amount (positive number)
note: Optional note for this contribution
Returns:
Updated goal details
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| amount | Yes | ||
| goal_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is a non-read-only, non-idempotent mutation. The description adds that the tool 'Returns updated goal details' and makes the mutation explicit with 'Add a contribution,' but it does not explain side effects such as whether a transaction is created or account balances change.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-structured docstring with one straightforward purpose sentence, a compact Args list, and a clear Returns line. No filler; the operation is front-loaded and each section earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple three-parameter tool with no output schema, the description provides all invocation parameters and a return-value summary. It could be more specific about what 'updated goal details' contains, but the definition is sufficient for selecting and calling the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema description coverage is 0%, the Args section defines every parameter with real meaning: goal_id identifies the goal, amount is a positive contribution amount, and note is optional. This goes well beyond the bare schema titles and types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('Add a contribution') and target ('financial goal'), which clearly differentiates it from sibling goal tools like create_goal, update_goal, and delete_goal. The description is unambiguous about the resource being acted on.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given for when to use this tool versus alternatives such as create_goal, update_goal, or link_goal_plan. The description implies contributing to an existing goal but does not state selection criteria, exclusions, or related tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_accountAccounts: Create accountAInspect
Create a new financial account.
Args:
name: Account name (e.g., "Chase Checking", "Visa Credit Card")
account_type: One of: checking, savings, cash, credit_card, loan, mortgage,
student_loan, heloc, arm, investment, other_asset, other_debt
balance: Opening balance (default: 0.00, can be negative for debt accounts)
institution: Bank or financial institution name
institution_login_url: Where the user logs in at this institution
(https only, e.g. "https://chase.com"; bare domains are
normalized to https). Optional; intended for a URL the user
provided, not a guessed or invented one. Plaid-linked accounts
get one automatically.
currency: Currency code. App is locked to USD pre-launch — any other
value is rejected. (See docs/CURRENCY_RESTORE_TODO.md to re-enable
multi-currency.)
interest_rate: Annual interest rate as percentage (0-100). For HELOC, this
is the draw-phase rate. For ARM, this is the initial fixed rate.
credit_limit: Credit limit (for credit card accounts)
minimum_payment: Minimum monthly payment (for debt accounts)
term_months: Original loan term in months (e.g. 360 for a 30-year mortgage)
loan_start_date: Loan origination date (YYYY-MM-DD)
loan_end_date: Loan payoff target date (YYYY-MM-DD)
original_balance: Original loan/debt amount when first taken out
monthly_escrow_tax: Monthly property-tax escrow (mortgage only)
monthly_escrow_insurance: Monthly insurance escrow (mortgage only)
draw_period_months: Number of months in the HELOC draw-down phase
draw_amount_per_month: Monthly draw amount during HELOC draw-down phase
rate_2: 2nd-term annual interest rate as percentage — HELOC repayment
rate or ARM adjusted rate
term_2_months: 2nd-term duration in months — HELOC repayment term or
ARM remaining term after the adjustment
arm_initial_period_months: Fixed-rate period in months before ARM
adjustment (e.g. 60 for a 5/1 ARM)
include_in_debt_paydown: Whether this debt is part of the payoff
calculator and the dashboard payoff card (default True). Set
False for a debt you pay in full every month (e.g. a credit
card that carries no balance month to month) so it doesn't distort the
plan. Ignored for asset accounts; does NOT affect net worth,
total debt, or the balance sheet.
low_balance_alert_threshold: Per-account low-balance alert
threshold in dollars (checking/savings/cash only). Omit to use
the monitoring plan's global threshold.
mute_low_balance_alerts: Suppress low-balance alerts on this account
(daily brief + notifications). For accounts the user
intentionally keeps low. Default False.
Returns:
Created account details
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| rate_2 | No | ||
| balance | No | ||
| currency | No | USD | |
| institution | No | ||
| term_months | No | ||
| account_type | Yes | ||
| credit_limit | No | ||
| interest_rate | No | ||
| loan_end_date | No | ||
| term_2_months | No | ||
| loan_start_date | No | ||
| minimum_payment | No | ||
| original_balance | No | ||
| draw_period_months | No | ||
| monthly_escrow_tax | No | ||
| draw_amount_per_month | No | ||
| institution_login_url | No | ||
| include_in_debt_paydown | No | ||
| mute_low_balance_alerts | No | ||
| monthly_escrow_insurance | No | ||
| arm_initial_period_months | No | ||
| low_balance_alert_threshold | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=false; the description does not contradict these and adds genuinely useful behavior: currency is hard-locked to USD and other values are rejected, bare domains in institution_login_url are normalized to https, Plaid-linked accounts get a login URL automatically, and include_in_debt_paydown explicitly does NOT affect net worth, total debt, or the balance sheet. It stops short of explaining duplicate-account behavior or what happens on partial failure, which would be relevant given idempotentHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Long, but each parameter entry earns its place given 23 undocumented-in-schema fields, and the one-line purpose statement is front-loaded. The trailing 'Returns: Created account details' is near-vacuous but minor. Some entries (currency's docs-file reference) are slightly noisy but still informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 23-parameter creation tool with no output schema and only hint-level annotations, this is close to complete: it covers defaults, valid ranges, enum values, and cross-parameter applicability rules. Remaining gaps are return-shape detail (addressed minimally), duplicate/idempotency behavior, and error conditions like invalid account_type values, which matter given idempotentHint=false.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there are 23 parameters, so the description carries the full burden — and it largely does. It enumerates all 11 valid account_type values (the schema has no enum), gives examples and units for name, balance, interest_rate (0-100 annual %), term_months (360 = 30-year mortgage), dates in YYYY-MM-DD, and scopes conditional parameters (escrow fields = mortgage only, draw fields = HELOC only, low-balance fields = checking/savings/cash only). This is meaning far beyond what the bare typed schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a new financial account'), which distinguishes it from sibling creators like create_transaction, create_goal, or create_hard_asset by resource type. It does not explicitly name those alternatives, but the resource noun plus the account_type enumeration makes the scope unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus import_accounts, create_hard_asset, or list_accounts-then-act workflows. There is also no mention of prerequisites (e.g. whether the account must be linked to a Plaid item afterward). The description moves straight into parameter documentation with no usage context at all.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_budgetBudgets: Create budgetAInspect
Create a new budget for an expense category.
Args:
category_name: Name of the expense category to budget (case-insensitive)
amount: Monthly/yearly budget amount (minimum: 0.01)
period: Budget period - "monthly" or "yearly" (default: monthly)
start_date: Start date in YYYY-MM-DD format (defaults to 1st of current month)
rollover: Carry unspent budget to next period (monthly budgets only, default: False)
name: Optional display name. Defaults to the category name when blank. Useful
when a user wants multiple visibly-distinct budgets without duplicating
the category itself (Phase 92 #K1).
Returns:
Created budget details
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| amount | Yes | ||
| period | No | monthly | |
| rollover | No | ||
| start_date | No | ||
| category_name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only indicate non-readOnly, non-idempotent, non-destructive, and openWorldHint. The description adds the default behavior for period and start_date, the limitation that rollover is only for monthly budgets, and the behavior of name defaulting to category name. These are useful behavioral details that go beyond the annotations, though it does not explicitly state side effects like potential creation of a category if not exists, which is why it's not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a clear Args list and Returns section. It is slightly verbose due to the detailed explanation of each parameter, but this is justified given the absence of schema descriptions. The information is front-loaded with the core purpose before details, and every line adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers all six parameters, their defaults, constraints, and edge cases (rollover limited to monthly). It also explains the return value meaning. There is no output schema, but the description does a good job. It is slightly lacking in that it does not mention whether the tool checks for existing budgets or potential conflicts, but overall it is fairly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Even though schema description coverage is 0%, the description explicitly defines each parameter in the Args list, including constraints (minimum amount, valid period options, date format, and the special behavior of name). This adds meaning beyond the schema, which only provides types and defaults. It compensates well for the lack of schema descriptions, though not for nested semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'create' and the resource 'budget for an expense category', which is specific and actionable. It is distinct from siblings like create_category and create_goal, and the inclusion of 'create a budget' makes it unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides detailed parameter semantics including the purpose of the optional display name, which helps distinguish when to use this tool over creating a new category. It also implies the use case for multiple visibly-distinct budgets, effectively explaining when a user would want to use this tool instead of creating a category.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_categoryCategories: Create categoryAInspect
Create a new custom category.
Args:
name: Category name (max 50 chars)
category_type: "income" or "expense" (default: expense)
Returns:
Created category details
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| category_type | No | expense |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is not read-only, and the description confirms a creation side effect. It adds the 'custom category' framing and notes that it returns the created details, but it does not disclose behavior such as duplicate handling, uniqueness constraints, or whether creation can fail. This is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-organized with Args and Returns sections. Every sentence contributes necessary information: the action, the parameter constraints, and the return behavior. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter create tool, the description covers the essential inputs and return value. It could be slightly more complete by specifying the shape of the created category details or error conditions, especially since no output schema exists, but it is sufficiently complete for correct invocation in most cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter semantics. It does so fully: name is limited to 50 characters, category_type is restricted to 'income' or 'expense', and the default is 'expense'. This adds meaningful meaning beyond the raw schema, which only lists string types and a default.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'Create a new custom category.' This clearly distinguishes it from sibling tools like update_category, delete_category, and list_categories, and from other create_* tools such as create_tag or create_transaction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives, such as when a custom category is appropriate versus using existing categories or creating a tag. There are no stated exclusions, prerequisites, or decision heuristics, leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_debt_payment_ruleDebt Payments: Create debt-payment ruleAInspect
Create a rule that auto-matches withdrawals on an asset account to a debt account.
When a transaction on `source_account_id` has a description containing
the `description_pattern` (case-insensitive substring), the system
creates a pending debt payment that the user can confirm or reject.
After the configured threshold of confirmations, matches auto-apply.
Args:
source_account_id: Asset account the payment leaves from (e.g. checking)
debt_account_id: Debt account the payment reduces (e.g. mortgage)
description_pattern: Substring to match on the source transaction (<=200 chars)
Returns:
Created or reactivated rule details
| Name | Required | Description | Default |
|---|---|---|---|
| debt_account_id | Yes | ||
| source_account_id | Yes | ||
| description_pattern | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Reveals non-obvious behavior beyond the readOnlyHint=false annotation: matched transactions create pending debt payments that users confirm or reject, and matches auto-apply only after a configured threshold. This gives the agent important expectations about side effects and workflow.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, followed by a compact Args block and a Returns line. The length is justified because the deferred matching workflow is genuinely complex; no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool with no output schema, it covers the key mechanics needed to call correctly: source account, debt account, matching pattern, pending-payment lifecycle, and auto-apply behavior. It leaves the threshold configuration and duplicate/reactivation details somewhat implicit, but the provided guidance is sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description fully compensates by explaining all three parameters with practical examples (checking, mortgage), the substring match semantics, and the 200-character limit. Each parameter gains real meaning beyond its schema type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Create a rule that auto-matches withdrawals on an asset account to a debt account.' It goes beyond the title by explaining the matching behavior and differentiates this creation tool from update/delete/list rule siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for when to use the tool: when the user wants automatic matching of asset-account withdrawals to debt accounts via a description pattern. It does not explicitly name alternatives or exclusions, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_family_groupFamily: Create family groupAInspect
Create a new family group (user becomes owner). Requires Premium.
Args:
name: Group name
Returns:
Created group details
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is a mutation (readOnlyHint=false), so the description adds value by disclosing the Premium requirement and that the user becomes owner. There is no contradiction with annotations, and these extra details go beyond the structured metadata.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is succinct, front-loaded with the core action and requirement, and structured with clear Args/Returns sections. Every sentence serves a purpose with no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple creation tool with one parameter, the description covers the essential prerequisites and ownership outcome, but the return value is vaguely described as 'Created group details' with no output schema to fill the gap. This leaves some uncertainty about what the agent can expect, so a 3 is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the 'name' parameter. It provides 'name: Group name', which clarifies the purpose but adds little beyond the schema's title. Given the single simple parameter, this minimal addition is adequate, earning a 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action 'Create a new family group' and specifies the effect 'user becomes owner', distinguishing it from invitation-based siblings like accept_family_invitation or decline_family_invitation. The verb-resource pair is explicit and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool—creating a new family group—and notes the Premium requirement, but does not explicitly contrast with alternatives or state when not to use it. This is clear context without exclusions, fitting a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_forum_postForum: Create forum postAInspect
Create a new forum post.
Args:
category_slug: Slug of the category to post in
title: Post title (3-200 chars)
content: Raw markdown content (10-10000 chars)
Returns:
``{"success": True, "post": {...}}`` on success, or
``{"error": "..."}`` (unknown category, validation, or a staff-only
category for a non-staff user).
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| content | Yes | ||
| category_slug | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It discloses behavior beyond the annotations: it specifies that content is raw markdown, enforces character limits, returns a structured success/error dict, and mentions the staff-only category restriction. Since the annotations only contain false hints and offer no meaningful safety context, the description carries the burden and does so well, though it does not mention authentication requirements or post visibility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The docstring is well-structured with a one-sentence purpose, a clear Args list, and a Returns section. It is information-dense without being bloated; every sentence either defines a parameter or explains expected outcomes. The layout front-loads the core action and then provides supporting details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description appropriately includes both success and error return formats. It covers validation and staff-only edge cases, which helps the agent handle failures. It does not mention how to obtain a valid category_slug, making the tool slightly less self-contained, though the error for unknown categories provides a signal.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates by explaining each parameter: category_slug is a 'Slug of the category to post in', title has a 3-200 char limit, and content is 'Raw markdown content' with a 10-10000 char limit. This adds concrete meaning and usage constraints that the input schema alone does not provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Create a new forum post.' This clearly distinguishes it from sibling tools like create_forum_reply, which creates a reply. The title 'Forum: Create forum post' reinforces the resource without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by explaining what the tool does and what arguments are required, but it never explicitly states when to prefer this over create_forum_reply or how to discover valid category_slug values. The mention of 'staff-only category' errors hints at one exclusion condition, but there is no clear when-to-use or when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_forum_replyForum: Create forum replyAInspect
Reply to a forum post.
Args:
post_id: The parent post ID
content: Raw markdown content (5-10000 chars)
Returns:
``{"success": True, "reply": {...}}`` on success, or
``{"error": "..."}`` (post not found, locked, or validation).
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | ||
| post_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, so the write nature is known. The description adds minimal behavioral context by noting in Returns that the post might be locked, which hints at a potential failure condition. However, it does not disclose other side effects like permission requirements or rate limits, so it adds limited value beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured as a docstring with a clear one-line purpose, an Args section, and a Returns section. It front-loads the purpose and wastes no words. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there is no output schema, the description explicitly documents the return format and error cases. The two parameters are thoroughly explained. It lacks only minor context like prerequisites (e.g., the post must exist) but covers those via error messages. Overall, it is sufficiently complete for a simple creation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description fully compensates by explaining each parameter: post_id is 'the parent post ID' and content is 'Raw markdown content (5-10000 chars)'. This adds essential meaning beyond the schema's bare type definitions, including a character constraint.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the action 'Reply to a forum post' with a specific verb and resource. It clearly distinguishes from siblings like create_forum_post and update_forum_reply by the nature of the operation. No ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives such as update_forum_reply or delete_forum_reply. The description simply states what it does without contextual conditions or exclusions. An agent must infer usage from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_goalGoals: Create goalAInspect
Create a new financial goal.
Args:
name: Goal name (e.g., "Emergency Fund", "Vacation to Italy")
target_amount: Target amount to save (minimum: 0.01)
goal_type: One of: savings, debt_payoff, emergency_fund, custom (default: savings)
target_date: Target date in YYYY-MM-DD format (optional)
current_amount: Starting amount already saved (default: 0, ignored if linked_account)
linked_account_id: Link to an existing account for auto-tracking (optional)
color: Hex color code for the goal (default: #007bff)
notes: Optional notes about the goal
priority: Priority level (0 = highest, default: 0)
vng_goal_id: Optional Visions & Goals plan id to link (the
qualitative clarify-and-plan companion at visionsandgoals.com)
Returns:
Created goal details with progress
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| color | No | #007bff | |
| notes | No | ||
| priority | No | ||
| goal_type | No | savings | |
| target_date | No | ||
| vng_goal_id | No | ||
| target_amount | Yes | ||
| current_amount | No | ||
| linked_account_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral details: it returns 'Created goal details with progress' and notes that current_amount is ignored if linked_account is provided, which is additional context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the purpose, then uses a structured Args/Returns format. Each line earns its place for the 10 parameters, though it is somewhat lengthy and repeats type information that could be inferred from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 10 parameters and no output schema, the description covers the return value and all parameter semantics. It lacks details on error handling or explicit sibling differentiation, but for a create operation it provides enough to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must document all 10 parameters. It does so thoroughly: examples for name, minimum for target_amount, enum-like list for goal_type, date format for target_date, default and ignore condition for current_amount, hex format for color, and priority semantics. This fully compensates for the absent schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource: 'Create a new financial goal.' This is clear and direct. However, it does not explicitly differentiate from sibling tools like create_goal_plan, list_goals, or update_goal, leaving minor ambiguity about when this exact tool is preferred.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, nor any prerequisites such as required permissions or account state. The description only says what it creates, not when it is appropriate to call it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_goal_planGoals: Create goal planAInspect
Create a goal PLAN on Visions & Goals — the qualitative
clarify-and-plan side (SMART criteria, action steps). Suited to a
vague aspiration that needs shaping ("I want to get fit", "someday
buy a cabin"). The money-tracking side is a separate zoninga goal;
the two can be linked.
Args:
title: The goal statement (required, ≤300 chars).
description: Longer context for the plan.
timeframe: One of "short", "medium", "long" (default medium).
target_date: YYYY-MM-DD (optional).
smart_specific: Answer to the SMART "Specific" prompt (optional).
smart_measurable: SMART "Measurable" answer (optional).
smart_achievable: SMART "Achievable" answer (optional).
smart_relevant: SMART "Relevant" answer (optional).
smart_timebound: SMART "Time-bound" answer (optional).
zoninga_goal_id: An existing zoninga goal id to link the new
plan to (sets that goal's ``vng_goal_id``).
Returns:
{"goal": {id, title, ..., edit_url}}. ``edit_url`` is the page
on Visions & Goals where the plan can be refined.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| timeframe | No | medium | |
| description | No | ||
| target_date | No | ||
| smart_relevant | No | ||
| smart_specific | No | ||
| smart_timebound | No | ||
| zoninga_goal_id | No | ||
| smart_achievable | No | ||
| smart_measurable | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (non-readonly, non-idempotent, non-destructive, open world), so the bar is lower. The description still adds real behavioral context beyond them: the title length cap, the side effect of setting the linked goal's vng_goal_id, and the returned edit_url.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose followed by organized Args/Returns sections; the docstring formatting ('<' char cap, backtick field names) adds slight noise but each line carries information. No filler sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter mutation with no output schema, the description documents the return payload (goal id, edit_url) and every parameter's semantics, plus the linkage side effect. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the full burden and does: each of the 10 parameters is explained, including the timeframe enum values ('short','medium','long') that the schema leaves as a bare string and the date format for target_date. This is meaningful meaning beyond the structured fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create a goal PLAN on Visions & Goals') and immediately scopes it to the qualitative clarify-and-plan side, explicitly separating it from the money-tracking 'zoninga goal'. An agent can distinguish it from create_goal without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear usage condition ('Suited to a vague aspiration that needs shaping') with two concrete examples, and names the alternative domain (the money-tracking side) that is a separate goal. It stops short of explicitly naming the sibling tool to call instead, but the when-to-use guidance is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_hard_assetEstate: Create hard assetAInspect
Create a hard asset (real estate, vehicle, valuable) for net worth tracking.
Args:
name: Asset name (e.g., '123 Main St', '2022 Toyota Camry')
category: One of: real_estate, vehicle, jewelry, art, electronics, furniture, other
current_value: Current estimated value
purchase_price: Original purchase price (optional)
purchase_date: Purchase date in YYYY-MM-DD format (optional)
description: Description or notes (optional)
location: Physical location (optional)
linked_account_id: Associated debt account ID, e.g., mortgage (optional)
Returns:
Created hard asset details.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| category | Yes | ||
| location | No | ||
| description | No | ||
| current_value | Yes | ||
| purchase_date | No | ||
| purchase_price | No | ||
| linked_account_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish that this is a mutating operation (readOnlyHint=false, idempotentHint=false). The description adds the purpose of net worth tracking and states the return value, but it does not disclose edge behaviors such as duplicate handling, what happens to linked_account_id, or whether creating the asset immediately affects net worth calculations. This is acceptable but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose, then uses a clean Args block to document all parameters, followed by a Returns line. There is no fluff or repetition of schema structure beyond what is needed. Each line earns its place in supporting tool invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a create operation with 8 parameters and no output schema, this description is largely complete: it covers the purpose, semantics for every parameter, optionality, and a minimal return summary. It lacks detail on error behavior, required-field emphasis, and the exact shape of the created asset response, but these gaps do not block correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries full responsibility for parameter meaning. It does so thoroughly: it provides example name values, enumerates the allowed category values, specifies YYYY-MM-DD for purchase_date, marks optional parameters, and explains linked_account_id as an associated debt account such as a mortgage. This exceeds baseline compensation for schema gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Create a hard asset (real estate, vehicle, valuable) for net worth tracking.' It names the exact domain and the purpose, distinguishing it from other create_* tools such as create_account or create_category. The title 'Estate: Create hard asset' further reinforces the scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states a clear use context: creating a hard asset for net worth tracking. This implicitly tells an agent when to use the tool versus account or category creation tools. It does not explicitly name alternatives or provide a when-not-to-use statement, but the purpose clause is specific enough for correct selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_milestoneMilestones: Create milestoneAInspect
Create a milestone checkpoint within a goal.
Args:
goal_id: The parent goal ID
name: Short milestone name ("25% saved", "First $1k")
target_amount: Amount at which the milestone is reached (> 0 and <= goal target)
celebration_message: Optional message shown when reached
Returns:
Created milestone details
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| goal_id | Yes | ||
| target_amount | Yes | ||
| celebration_message | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only, non-destructive, non-idempotent write, so the safety profile is covered. The description adds a genuine behavioral constraint beyond the schema: target_amount must be > 0 and <= the goal target, which is the kind of validation detail the schema lacks.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the one-line purpose before the Args/Returns sections, and every line carries information. The docstring scaffolding (Args/Returns headers) adds minor verbosity but no wasted sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-idempotent mutation with no output schema, the definition covers purpose, all four parameters, the target_amount constraint, and confirms a created-milestone return. It omits side effects (e.g., whether reaching a milestone triggers the celebration notification) and error behavior, which are the remaining gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full parameter burden and largely does so: it explains goal_id as the parent goal, gives concrete name examples, constrains target_amount, and clarifies celebration_message is optional and shown on completion. It does not specify units/currency for target_amount, which is the only meaningful gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a milestone checkpoint within a goal'), which is distinctly more specific than the sibling list/create/update/delete_milestone names alone. It does not explicitly name sibling tools, so it stops short of a 5, but the scoping to a parent goal makes the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance and no mention of alternatives such as update_milestone or create_goal_plan. The requirement for an existing parent goal is only implied by the goal_id argument, not stated as a prerequisite.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_recurring_transactionRecurring Transactions: Create recurring templateAInspect
Create a recurring transaction template.
Args:
account_id: Source account
amount: Positive decimal in major units
description: Free-text description
frequency: One of daily, weekly, bi_weekly, monthly, yearly
start_date: ISO date (YYYY-MM-DD); first instance will spawn on this date
transaction_type: deposit | withdrawal | transfer (default: withdrawal)
category_id: Optional Category id
end_date: Optional ISO date when the recurrence stops
transfer_to_account_id: Required when transaction_type=transfer
Returns:
The created template (or ``{"error": ...}`` on validation failure).
| Name | Required | Description | Default |
|---|---|---|---|
| amount | Yes | ||
| end_date | No | ||
| frequency | Yes | ||
| account_id | Yes | ||
| start_date | Yes | ||
| category_id | No | ||
| description | Yes | ||
| transaction_type | No | withdrawal | |
| transfer_to_account_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only indicate that this is not read-only, not idempotent, and not destructive in the usual sense. The description adds meaningful behavioral context: the first instance spawns on start_date, validation failures return an error object, and transfer_to_account_id is conditionally required. This goes beyond what the annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a well-organized docstring with a one-sentence summary, a compact Args list, and a Returns line. Every line earns its place; there is no filler or redundancy, and the most important scoping statement is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter creation tool with no output schema and no parameter descriptions in the schema, the description is complete: it documents all required inputs, key defaults, conditional constraints, and the expected return shape including the error case. An agent has enough information to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden, and it succeeds. It explains every parameter: amount is a positive decimal in major units, frequency has an explicit enum, start_date has ISO format and spawn behavior, transaction_type has values and a default, and transfer_to_account_id is conditionally required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Create a recurring transaction template.' This clearly distinguishes it from the sibling create_transaction by emphasizing the recurring-template nature, so an agent can tell them apart without inspecting the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'recurring' implies this is for templated, repeating transactions rather than one-time ones, but the description never explicitly says when to use this tool versus create_transaction or update_recurring_transaction. Usage context is implied rather than stated, and no alternatives or exclusions are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_reminderAgent: Create scheduled reminderAInspect
Schedule a reminder for the user. It appears in their morning brief on
matching days ("Reminder: <text>"), and in the brief email when they
have daily-brief email delivery turned on. This is the right tool for
asks like "remind me every other Friday to pay child support".
Args:
text: What to remind them of (max 200 chars), e.g. "pay child support".
frequency: once / weekly / biweekly / monthly / yearly. weekly and
biweekly repeat on start_date's weekday; monthly repeats on its
day-of-month (clamped in short months); yearly repeats on its
month + day (e.g. a birthday).
start_date: First occurrence, YYYY-MM-DD.
end_date: Optional last day (YYYY-MM-DD) after which it stops.
Returns:
{"success": True, "reminder": {...}, "delivery_note"?: str}. Relay
delivery_note to the user when present — it means the brief cadence
won't actually deliver the reminder (plan inactive / cadence off or
non-daily).
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| end_date | No | ||
| frequency | Yes | ||
| start_date | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are minimal (readOnlyHint=false, destructiveHint=false, etc.), so the description carries the behavioral burden. It discloses side effects (appears in morning brief, email if enabled), the delivery_note caveat (won't actually deliver if cadence off), and date clamping semantics—all beyond annotation scope and consistent with them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is somewhat long but every sentence adds functional value: it opens with the core purpose, gives an example, then details each parameter and the return contract. Structure with Args/Returns aids scanning, though it could be slightly tightened without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with non-trivial recurrence rules and a special delivery_note return case, the description covers everything needed for correct invocation: parameter constraints, recurrence behavior, return format, and a critical user-facing relay instruction. No output schema exists, but the return contract is fully described.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must fully document parameters. It explains text max length, frequency values with detailed recurrence rules (weekly on start_date's weekday, monthly day clamping, yearly month+day), start_date format, and end_date optionality—far exceeding schema titles.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Schedule'), a specific resource ('a reminder'), and concrete behavior ('appears in their morning brief...'), with an example ask. This distinguishes it from siblings like delete_reminder and list_reminders without needing to open their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly provides a canonical usage example ('remind me every other Friday to pay child support') and explains where the reminder appears, which guides when to use it. However, it does not name alternatives or state when not to use it, so it stops short of the 'explicit when-not' bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_scenarioScenarios: Create a what-if scenarioAInspect
Create a what-if scenario and return its projected impact.
A scenario bundles one or more hypothetical changes (buy a car, take a
mortgage, open a credit card, sell the house). It auto-saves for 60
days; committing it later turns its add items into real accounts.
Each item is a dict:
- add_debt: {action:"add_debt", account_type (credit_card / loan /
mortgage / student_loan / heloc / arm / other_debt), name, balance,
interest_rate, minimum_payment, term_months; for HELOC/ARM also the
full two-phase set rate_2 + term_2_months + draw_period_months or
arm_initial_period_months}
- add_asset: {action:"add_asset", account_type (cash / checking /
savings / investment / other_asset), name, balance}
- remove: {action:"remove", target_account_id} — sells/drops an
existing account from the projection.
Args:
name: Short label for the scenario (e.g. "Buy a 2026 Honda").
items: List of change dicts (see shapes above).
description: Optional notes.
strategy: Paydown strategy for the projection (default highest_rate
/ avalanche).
extra_monthly: Extra monthly payment applied to the projection.
Returns:
{scenario: {...}, analysis: {...}} or {"error": "..."}.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| items | Yes | ||
| strategy | No | highest_rate | |
| description | No | ||
| extra_monthly | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), so the bar is lower, yet the description adds real behavioral context: the 60-day auto-save window, temporary nature, and that committing later materializes add items as real accounts. It does not mention auth/permission requirements, but the lifecycle disclosure is valuable beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose, then structured Args/Returns sections and a nested item-shape block. It is long but the length is justified by the genuinely complex nested item contract; the indentation is slightly scattered but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description responsibly indicates the return shape ({scenario, analysis} or {error}) and compensates for the 0% schema coverage with full parameter documentation. The internal structure of 'analysis' is left unspecified, which is a minor gap but not needed to invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the items parameter is a formless object with additionalProperties=true, so the description carries the entire burden. It fully documents the item dict shapes (add_debt with enumerated account_type values and HELOC/ARM two-phase fields, add_asset, remove) plus name, description, strategy default (highest_rate/avalanche), and extra_monthly, adding substantial meaning absent from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a what-if scenario and return its projected impact') and clarifies scope with concrete examples of what a scenario bundles (buy a car, take a mortgage). This clearly distinguishes it from siblings like commit_scenario, update_scenario, and list_scenarios without needing their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context for when to use it: a hypothetical/what-if bundling of changes that 'auto-saves for 60 days; committing it later turns its add items into real accounts.' This implies the create-vs-commit workflow, though it never explicitly names an alternative tool or states when NOT to use it, so it falls short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_spending_challengeGamification: Create spending challengeAInspect
Create a personal spending challenge.
Args:
name: Challenge name (e.g., "No Coffee November")
category_name: Category to track (e.g., "Dining", "Entertainment")
metric: "spending_limit" (stay under target) or "spending_avoidance" (spend $0)
target_amount: Target amount (used for spending_limit; set 0 for avoidance)
duration_days: Duration in days (7 or 30)
Returns:
Created challenge details
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| metric | Yes | ||
| category_name | Yes | ||
| duration_days | No | ||
| target_amount | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is a write operation (readOnlyHint=false, destructiveHint=false, idempotentHint=false). The description adds the metric semantics (spending_limit vs spending_avoidance) and the target_amount behavior (set 0 for avoidance), which is useful. However, it doesn't disclose what happens on creation (e.g., whether it's immediately active, whether it affects gamification points, or whether duplicate challenges are allowed).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured with a clear Args list. Every line adds value, and the parameter explanations are front-loaded. It could be slightly more concise by removing the Returns line, but it's not bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a creation tool with no output schema, the description covers the parameters well but omits behavioral context like whether the challenge starts immediately, whether there are constraints on active challenges, or what the response contains beyond 'Created challenge details'. The sibling list shows related tools (list_challenges, join_challenge) but the description doesn't connect to them.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden. It explains the meaning of each parameter, including the relationship between metric and target_amount, and provides examples for name and category_name. It also clarifies that duration_days is 7 or 30, which the schema does not. This is strong compensation for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb ('Create') and resource ('personal spending challenge'), and the title adds the 'Gamification' context. It distinguishes itself from sibling tools like list_challenges and join_challenge by focusing on creation, though it doesn't explicitly name those siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (when creating a spending challenge) but provides no explicit guidance on when not to use it or alternatives. Siblings like join_challenge and list_challenges exist, but the description doesn't mention them or explain the difference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_support_ticketSupport: Create support ticketAInspect
File a new support ticket on the user's behalf.
Writes to Zoninga's own support desk (no third-party service). Staff
handle the ticket on the web support hub and the user is notified of
replies. API docs: https://zoninga.com/ai/#support-tickets
Args:
subject: Short summary (1-200 chars).
body: The full message describing the issue (1-5000 chars).
category: One of ``bug``, ``feature``, ``billing``, ``account``,
``other`` (default ``other`` — staff re-triage anyway).
Priority is NOT settable here — staff triage urgency.
Returns:
``{"success": True, "ticket": {...}}`` (the new ticket, with its
initial message counted), or ``{"error": "..."}`` on validation
failure.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | ||
| subject | Yes | ||
| category | No | other |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds real context beyond that: it writes to Zoninga's own desk (not a third party), staff triage on the web hub, the user is notified of replies, priority is deliberately not settable, and it gives the return shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in one sentence, then Args/Returns are cleanly structured with no filler. Slightly longer than strictly necessary (docs URL, triage rationale) but every line adds usable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param mutation tool with no output schema, the description fully documents inputs, defaults, return shape (success/error), and downstream behavior. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden and does: it gives length bounds for subject (1-200) and body (1-5000), enumerates the allowed category values, states the default (other) with rationale, and clarifies that priority is intentionally absent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("File a new support ticket on the user's behalf") and the word "new" plus "on the user's behalf" cleanly separates it from reply_to_support_ticket, get_support_ticket and list_support_tickets. An agent can identify the action without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the context (creating a fresh ticket for the user) but never states when to use this versus reply_to_support_ticket or when not to file one. The only routing signal is the note that priority is not settable here, which is helpful but not a when-to-use rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_tagTags: Create tagBInspect
Create a tag.
Args:
name: Tag name (max 30 chars), unique per user.
color: Hex color code like ``#007bff`` (default: gray)
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| color | No | #6c757d |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false and idempotentHint=false, so the safety profile is covered structurally. The description adds one genuinely useful behavioral fact: names are unique per user. It does not say what happens on a duplicate name (error vs. upsert), which matters given idempotentHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short lines, purpose first, then arguments with constraints. No filler, no restatement of the tool name beyond the imperative verb.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter creation tool with no output schema and no nesting, the argument guidance is adequate, but the description never says what is returned (e.g. the new tag's id) or how name conflicts surface. With no output schema, that gap falls on the description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden and largely does: it documents the 30-char cap and per-user uniqueness for 'name' and the hex format (``#007bff``) for 'color'. Only the color default is left to the schema's default field, and validation behavior for malformed hex is unstated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create a tag'), which distinguishes it from update_tag, delete_tag and list_tags in the sibling set. It does not, however, say what a tag is for or how it differs from the similarly-shaped create_category or create_milestone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to create a tag versus when to reuse an existing one, no mention of list_tags for discovery or update_tag for modification, and no prerequisites or exclusions. The agent must infer all routing on its own.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
create_transactionTransactions: Create transactionAInspect
Create a new transaction.
Args:
account_id: The account ID for this transaction
amount: Transaction amount (positive number)
description: Description of the transaction
transaction_type: 'deposit', 'withdrawal', or 'transfer'.
On asset accounts (checking, savings, cash, investment):
- 'deposit' = money in (pair with an income category like "Salary")
- 'withdrawal' = money out (pair with an expense category like "Groceries")
On debt accounts (credit_card, loan, mortgage, heloc, arm, student_loan):
- 'deposit' = a charge/purchase that INCREASES the debt
(pair with an expense category like "Shopping")
- 'withdrawal' = a payment that DECREASES the debt
(pair with an income category like "Credit Card Payment")
Record a CC purchase as transaction_type='deposit' with an expense
category; record a CC bill payment as transaction_type='withdrawal'
with an income category. The category type is inverted on debt
accounts versus asset accounts — this is by design.
transaction_date: Date in YYYY-MM-DD format (defaults to current UTC date)
category_name: Category name (optional, case-insensitive)
transfer_to_account_id: Destination account ID (required for transfers)
confirm_future: If transaction_date is more than 30 days in the
future, the tool refuses and returns a `requires_confirmation`
error. Resubmit with confirm_future=True to accept — e.g.
scheduled rent a few months out. Defaults to False to catch
fat-fingered dates like "2035-01-01".
Returns:
Created transaction details
| Name | Required | Description | Default |
|---|---|---|---|
| amount | Yes | ||
| account_id | Yes | ||
| description | Yes | ||
| category_name | No | ||
| confirm_future | No | ||
| transaction_date | No | ||
| transaction_type | Yes | ||
| transfer_to_account_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes far beyond the sparse annotations by explaining the date default, the 30-day future-date guard, the requires_confirmation error, and the inverted category semantics on debt accounts. It also clearly states defaults for transaction_date, category_name, and confirm_future. This gives the agent accurate behavioral expectations before invoking the tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is somewhat long, but the length is justified by the complexity of transaction_type semantics and the confirmation behavior. It is well-structured as Args and Returns, with the core action front-loaded. A few explanatory sentences could be tightened, but nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex creation tool with no output schema, the description covers required arguments, optional arguments, defaults, transfer requirements, and a key error case. The only gap is that the return value is only described as 'Created transaction details' without specifying which fields the agent can expect, which matters more because no output schema exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry all parameter meaning, and it does. Every parameter is explained, with especially valuable semantic detail for transaction_type, including concrete deposit/withdrawal examples for asset and debt accounts. confirm_future's purpose and failure mode are also clearly tied to the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear, specific action: 'Create a new transaction.' It names the resource and is easily distinguished from siblings like create_recurring_transaction or import_transactions. The detailed type guidance further clarifies what kind of transaction entity this creates.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong in-tool usage guidance, such as how transaction_type behaves differently on asset versus debt accounts and that transfer_to_account_id is required for transfers. However, it never explicitly tells the agent when to choose this tool over related siblings like create_recurring_transaction or update_transaction. There is no alternative or exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
decline_family_invitationFamily: Decline family invitationAInspect
Decline a pending family invitation.
Args:
invitation_id: ID of the invitation to decline
Returns:
Confirmation.
| Name | Required | Description | Default |
|---|---|---|---|
| invitation_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is not read-only and not destructive. The description adds the constraint that the invitation must be pending and notes that the return value is a confirmation, which is useful but limited. It doesn't disclose side effects, permission requirements, or irreversibility, though the annotation set lowers the burden somewhat.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-organized, front-loading the action in the first sentence and then providing structured Args and Returns sections. Every line earns its place with no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter mutation tool with no output schema, the description covers the core action, the parameter, and the return type. It is sufficient for an agent to invoke the tool correctly, though adding a note about recipient versus sender roles would improve it further.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only an integer parameter name with no description, and schema description coverage is 0%. The description compensates by explicitly stating that invitation_id is 'ID of the invitation to decline,' giving the parameter clear semantic meaning beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Decline') and the resource ('a pending family invitation'), which is specific enough to understand the tool's core purpose. It doesn't explicitly differentiate from sibling tools like accept_family_invitation or revoke_family_invitation, but the verb itself makes the basic distinction obvious.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives such as accept_family_invitation, send_family_invitation, or revoke_family_invitation. The description only restates the action without offering decision criteria or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_accountAccounts: Delete accountADestructiveInspect
Delete (soft-delete) a financial account. The account is deactivated, not permanently removed.
Args:
account_id: The account ID to delete
Returns:
Confirmation of deletion
| Name | Required | Description | Default |
|---|---|---|---|
| account_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark destructiveHint=true, but the description adds meaningful context: the deletion is a soft-delete and the account is merely deactivated, not permanently removed. It also discloses that a confirmation is returned. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured with Description, Args, and Returns sections. Every sentence adds information, and the soft-delete clarification is front-loaded. No redundant or filler text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with destructive annotations and no output schema, the description covers the essential behavior: soft-delete, deactivation, and confirmation. The return specification is somewhat vague ('Confirmation of deletion'), but given the low complexity and annotations, the definition is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only the name and type for account_id (integer), and schema description coverage is 0%. The description's Args section states 'account_id: The account ID to delete,' which adds minimal semantic context but largely restates what the parameter name already implies. For a single simple ID, this is acceptable but not rich.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Delete (soft-delete) a financial account.' It clearly distinguishes the action from permanent deletion by explaining the account is deactivated, not permanently removed. The title 'Accounts: Delete account' reinforces the resource without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage through the tool name and the phrase 'Delete a financial account,' and it clarifies the soft-delete semantics. However, it does not explicitly name alternatives (e.g., update_account) or state when not to use this tool, so selection guidance is only implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_agent_chat_sessionAgent Chat: Delete agent chat sessionADestructiveInspect
Delete one of the user's AI assistant chat sessions, including
every message in it.
This is irreversible — the cron-driven 90-day retention sweep
does the same thing for old sessions, but this tool is for users
who want to clear a session immediately (e.g. they shared
context they no longer want stored).
Args:
session_id: The chat session's ID.
Returns:
``{"success": True, "deleted_session_id": ..., "deleted_messages": N}``
or ``{"error": "..."}``.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, but the description adds real substance: it stresses irreversibility, states exactly what is destroyed (the session plus every message), and explains how it relates to the automated retention sweep. For a destructive tool this is the right level of disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose and the irreversibility warning in the first two paragraphs, then uses compact Args/Returns sections. The Google-style docstring scaffolding is slightly heavier than needed for a one-parameter tool, but no sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the return contract ({'success': True, 'deleted_session_id': ..., 'deleted_messages': N} or {'error': ...}), covers the destructive semantics, and gives the motivating use case. An agent has everything needed to call it correctly and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single parameter, so the description carries the burden, but its 'Args' entry ('The chat session's ID') merely restates the schema title 'Session Id' and adds no format, type, or sourcing detail (e.g. obtainable from list_agent_chat_sessions). Minimum viable rather than genuinely informative.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete one of the user's AI assistant chat sessions') and immediately scopes it ('including every message in it'). This distinguishes it from the many other delete_* siblings such as delete_family_chat_message without requiring schema inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use context: for users who want to clear a session immediately, e.g. after sharing context they no longer want stored, versus waiting for the 90-day retention sweep. It does not name sibling tools like get_agent_chat_session or list_agent_chat_sessions as sources for the ID, so it stops short of full alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_budgetBudgets: Delete budgetBDestructiveInspect
Delete a budget (soft delete).
Args:
budget_id: The ID of the budget to delete
Returns:
Confirmation of deletion
| Name | Required | Description | Default |
|---|---|---|---|
| budget_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the operation as destructive, but the description adds the important nuance of a 'soft delete,' which is genuinely informative beyond the annotation. It also states that a confirmation of deletion is returned, giving a bit of behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The main purpose is front-loaded in the first sentence, and the rest is a compact Args/Returns structure. It is slightly redundant because the Args section repeats the schema, but it remains short and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter soft-delete operation with a destructive annotation, the description is minimally adequate: it names the resource, the parameter, and the return type. However, it leaves the meaning of 'soft delete' vague and gives only a generic confirmation as the return value, with no output schema to fill that gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter description, 'budget_id: The ID of the budget to delete,' essentially restates the schema property title 'Budget Id.' It adds no guidance on how to obtain the ID, its format, or any validation rules, so it does little to compensate for the schema description coverage of 0%.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'Delete a budget (soft delete).' This is clear and distinct from nearby siblings like update_budget or list_budgets, though it does not explicitly name or contrast any sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus alternatives such as update_budget, deactivating a budget, or other delete operations. No context, prerequisites, or exclusions are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_categoryCategories: Delete categoryADestructiveInspect
Delete a user-created category. System categories cannot be deleted.
Args:
category_id: The category ID to delete
Returns:
Confirmation of deletion
| Name | Required | Description | Default |
|---|---|---|---|
| category_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructiveHint=true and readOnlyHint=false, so the destructive nature is known. The description adds value beyond annotations by noting the system-category restriction and stating that a confirmation of deletion is returned. These are relevant behavioral details not present in the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the action, and the system-category restriction is clearly stated early. The Args and Returns sections are mostly redundant with the schema, but the overall size is appropriately brief and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter destructive action, the description covers the core operation, the key restriction, and the return behavior. It does not mention failure modes, cascading effects on transactions, or error handling, leaving some uncertainty for an agent invoking deletion.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the undocumented parameter, but it only restates 'category_id' as 'The category ID to delete' without adding real semantics. It does not explain how to obtain a valid user-created category ID, constraints, or why the ID might be invalid.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Delete a user-created category') and resource, and it explicitly scopes what can be deleted, distinguishing it from system categories. It also stands apart from siblings like create_category/update_category/list_categories. This is a clearly differentiated purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives useful usage context by specifying that system categories cannot be deleted, which implicitly tells the agent when not to call this tool. It does not name an alternative tool for renaming or hiding categories, but it provides a meaningful exclusion that helps route the call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_debt_payment_ruleDebt Payments: Delete debt-payment ruleADestructiveInspect
Deactivate a debt payment rule (soft delete).
Existing DebtPayment records created by the rule are retained for audit.
Args:
rule_id: ID of the rule to deactivate
Returns:
Confirmation of deactivation
| Name | Required | Description | Default |
|---|---|---|---|
| rule_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnly=false. The description adds meaningful context beyond this: the delete is soft, and existing DebtPayment records are retained for audit. This helps the agent communicate consequences accurately, though reversibility is not addressed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: purpose first, then the key behavioral consequence, followed by parameter and return details. Every sentence adds relevant information without filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter soft-delete tool, the description adequately covers the action, the retention behavior, and the return semantics. Additional details such as permission requirements or behavior on an already-deactivated rule would be nice but are not critical.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the schema only provides the parameter name 'rule_id'. The description compensates by explicitly defining it as 'ID of the rule to deactivate', giving the parameter its operational meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Deactivate') and resource ('debt payment rule'), and clarifies the semantic as a soft delete. It also distinguishes itself from a hard delete by noting audit records are retained.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The usage is implied: if you want to deactivate a rule, use this tool. However, there is no explicit guidance about when to prefer it over update_debt_payment_rule, or whether deactivated rules can be reactivated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_family_chat_messageFamily: Delete family chat messageADestructiveInspect
Delete a family chat message. Users can delete their own messages at
any time; group owners and admins can delete anyone's messages.
Args:
message_id: ID of the message to delete
Returns:
Confirmation on success, or error dict
| Name | Required | Description | Default |
|---|---|---|---|
| message_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the mutation semantics are covered structurally. The description adds value beyond that with the permission model governing whose messages may be deleted, plus a note that the call returns confirmation or an error dict. It does not discuss irreversibility or side effects on chat history, keeping it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences plus Args/Returns blocks; the purpose and permission rules are front-loaded where an agent will read them. The explicit 'Args:'/'Returns:' scaffolding is slightly heavier than needed for a single-parameter tool but wastes little space.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by stating the return shape ('Confirmation on success, or error dict'). Combined with the permission rules and annotations that already mark the operation destructive, an agent has everything needed to invoke this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only one parameter, an integer message_id, and 0% schema description coverage, the schema itself explains nothing. The description's 'Args: message_id: ID of the message to delete' is effectively a restatement of the parameter name and adds no format, sourcing, or validation detail beyond what the schema already implies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a specific verb and resource ('Delete a family chat message'), which is unambiguous against siblings such as edit_family_chat_message, send_family_chat_message, and list_family_chat_messages. It stops short of explicitly naming those alternatives, so it earns a clear 4 rather than a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description supplies the operative precondition for use: 'Users can delete their own messages at any time; group owners and admins can delete anyone's messages.' This tells the agent when the call is authorized, which is genuine routing-relevant context, though it never contrasts this tool with the related edit/send/list chat siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_forum_postForum: Delete forum postADestructiveInspect
Delete a forum post (author or staff). Replies cascade.
Args:
post_id: The post to delete.
Returns:
``{"success": True, "deleted_post_id": post_id}`` on success, or
``{"error": "..."}`` (not found, or not author/staff).
| Name | Required | Description | Default |
|---|---|---|---|
| post_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this as destructive, and the description adds meaningful behavioral detail beyond that: replies cascade, authorization is limited to author/staff, and errors are returned for not-found or unauthorized cases. The exact success and error return shapes are also disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, with the core behavior stated first, followed by minimal Args and Returns sections. Every sentence adds value and there is no redundant or promotional language.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one parameter, and the description covers the key requirements: what is deleted, who may delete, what happens to replies, and what response to expect on both success and failure. No output schema exists, but the return contract is explicitly documented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must clarify the parameter, and it does: 'post_id: The post to delete.' It also reinforces that the returned deleted_post_id is the same identifier, which is useful given the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Delete a forum post') on a specific resource, and clarifies authorization ('author or staff'). It also distinguishes posts from replies by noting that replies cascade, which separates it from delete_forum_reply.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: it is for posts, not replies, and only the author or staff can perform the deletion. It does not explicitly name alternatives like delete_forum_reply, but the resource scope and authorization condition are enough to guide correct usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_forum_replyForum: Delete forum replyADestructiveInspect
Delete a forum reply (author or staff).
Args:
reply_id: The reply to delete.
Returns:
``{"success": True, "deleted_reply_id": reply_id}`` on success, or
``{"error": "..."}`` (not found, or not author/staff).
| Name | Required | Description | Default |
|---|---|---|---|
| reply_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false. The description adds meaningful behavioral detail: the authorization requirement, the exact success return shape, and the error conditions (not found, or not author/staff). This goes beyond the structured annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, with the core action stated first. The Args/Returns formatting organizes the information efficiently with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter destructive action with no output schema, the description covers all essential aspects: target identification, authorization, success return value, and error cases. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry parameter meaning. 'reply_id: The reply to delete' clarifies that the integer identifies the target reply and is required. It is minimal but adequate for the single simple parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Delete a forum reply.' It also states the permission context (author or staff), which distinguishes it from related tools like update_forum_reply or delete_forum_post. The purpose is immediately unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool by noting that only the author or staff can delete. It does not explicitly name alternatives or exclusions, but the forum-reply scope and permission condition make selection straightforward among siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_goalGoals: Delete goalADestructiveInspect
Cancel a financial goal (soft delete).
Args:
goal_id: The ID of the goal to cancel
Returns:
Confirmation of cancellation
| Name | Required | Description | Default |
|---|---|---|---|
| goal_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnly=false, and the description adds the meaningful nuance that this is a soft delete, not a permanent hard deletion. It also specifies that a confirmation is returned. This adds context beyond the structured annotations, though it stops short of describing reversibility or impact on related data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly structured into purpose, argument, and return sections. There is no filler or redundant explanation, and the core action is front-loaded. Every element earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter soft-delete tool, the description gives the essential information: what it does, the parameter meaning, and the confirmation return. It lacks notes on post-cancellation visibility, reversibility, or effects on linked resources, but it is still nearly complete for invoking the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one required integer parameter with no description (0% coverage). The description compensates with an Args section stating that goal_id is 'The ID of the goal to cancel.' For a single, simple parameter this is sufficient, though it does not indicate where to obtain valid goal IDs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Cancel') and resource ('a financial goal'), and clarifies the mechanism ('soft delete'). This is easily distinguishable from the many sibling delete_* tools because it names the exact target resource and the action type.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, nor any exclusions. It does not mention related tools like update_goal, list_goals, or link_goal_plan, nor what happens to linked plans or shared goals when a goal is cancelled. Usage is implied only by the verb 'Cancel'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_hard_assetEstate: Delete hard assetADestructiveInspect
Delete a hard asset (soft-deactivate).
Args:
asset_id: ID of the hard asset to delete
Returns:
Confirmation message.
| Name | Required | Description | Default |
|---|---|---|---|
| asset_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide destructiveHint=true, so the description wisely adds the key nuance 'soft-deactivate,' indicating the deletion is not necessarily a permanent hard erase. It also discloses the return contract ('Confirmation message'), adding value beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: a one-line action statement followed by Args and Returns sections. There is no filler, and each section earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter mutation with no output schema, the description covers the action, the parameter, and the return value. However, it omits consequences such as reversibility, effects on related data, or whether the asset remains visible in any list, leaving the agent to infer what 'soft-deactivate' means operationally.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the Args line must carry the semantic load. It correctly identifies asset_id as 'ID of the hard asset to delete,' which is sufficient for a single required integer, but it adds little beyond the schema title and gives no source, format, or prerequisite details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb-resource operation: 'Delete a hard asset (soft-deactivate).' The parenthetical removes ambiguity about whether this is a permanent physical delete, and it clearly distinguishes the tool from sibling update_hard_asset and create_hard_asset.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use or when-not-to-use guidance, and no alternative tools are named. The purpose line implies the tool is for removing or deactivating a hard asset, but the agent is left to infer edge cases and exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_milestoneMilestones: Delete milestoneADestructiveInspect
Delete a milestone. (Milestones are hard-deleted — they're not user-facing records.)
Args:
milestone_id: The milestone ID to delete
Returns:
Confirmation of deletion
| Name | Required | Description | Default |
|---|---|---|---|
| milestone_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description reveals that milestones are hard-deleted and not user-facing records, which goes beyond the destructiveHint annotation by explaining the consequence. It also states the expected return (confirmation of deletion), adding useful behavioral context that the structured annotations do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded, with the core action in the first sentence. The Args and Returns sections are standard and minimal, containing no filler or redundant content. Every sentence serves a purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter delete operation with destructiveHint and idempotentHint annotations provided, the description covers the action, parameter, and return value. It does not address error cases or prerequisites, but these are not critical for a straightforward deletion tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only the parameter name and type with no description (0% coverage), so the description's clarification that milestone_id is 'the milestone ID to delete' is essential and directly compensates. It fully explains the single parameter's meaning and role.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states 'Delete a milestone' with a specific verb and resource, and the parenthetical clarifies hard-deletion, which distinguishes it from update_milestone or create_milestone. The intent is unambiguous and well-scoped.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives, no prerequisites, and no mention of exclusions. Usage is only implied by the name and the action 'Delete a milestone', which is insufficient for distinguishing appropriate invocation context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_notificationNotifications: Delete notificationADestructiveInspect
Permanently delete a notification from the user's inbox.
Deletion removes the notification from the user's history entirely
(e.g. dismissing an old Weekly Recap row); marking it read is the
non-destructive way to clear an unread badge.
Args:
notification_id: ID of the notification to delete.
Returns:
{"success": True, "deleted_id": ...} on success,
{"error": "..."} if the notification doesn't exist or isn't
owned by the caller.
| Name | Required | Description | Default |
|---|---|---|---|
| notification_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and readOnlyHint=false, so the safety profile is known. The description still adds value by stressing permanence ('removes the notification from the user's history entirely') and documenting failure semantics (error if the notification doesn't exist or isn't owned by the caller). It stops short of noting irreversibility/confirmation behavior explicitly, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then the sibling contrast, then Args/Returns. Every sentence earns its place; the only friction is the verbose docstring-style Args/Returns block for a single-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully documents the return shape (success/deleted_id on success, error string on missing/unauthorized). For a one-parameter destructive tool, deletion semantics, alternatives, and error cases are all covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and there is one integer parameter, so the description carries the burden. It gives only a bare restatement ('ID of the notification to delete'), adding no format, source, or constraint detail beyond the parameter name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Permanently delete a notification from the user's inbox') and immediately distinguishes the action from the sibling mark_notification_read by contrasting destructive deletion with non-destructive read-marking.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete use case ('dismissing an old Weekly Recap row') and an explicit alternative with its selecting condition: 'marking it read is the non-destructive way to clear an unread badge.' When-to-use and when-to-use-something-else are both covered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_payment_linkAgent: Delete payment linkADestructiveInspect
Delete one of the user's saved payment links permanently. After
deletion, payment discussions for that context fall back to the
account's bank link.
Args:
payment_link_id: The link's id.
Returns:
{"success": True, "deleted_id": ...} or an error when not found.
| Name | Required | Description | Default |
|---|---|---|---|
| payment_link_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false, so the safety profile is covered. The description adds genuinely new behavioral context: deletion is permanent and payment discussions for that context fall back to the account's bank link. It also discloses the return shape, though it doesn't cover error conditions in detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The body is short, front-loaded with the core action, and the consequence follows immediately. The Args/Returns scaffolding is somewhat boilerplate given the schema already defines the argument, but overall it is efficient and readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the return shape, and for a destructive mutation it discloses permanence and the downstream fallback behavior. Error handling is mentioned only vaguely ('an error when not found'), which is the main residual gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% for the single parameter, so the description must compensate. 'payment_link_id: The link's id' largely repeats the parameter name and adds no format, source, or lookup guidance (e.g., where to obtain the id), doing little to compensate for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete ... saved payment links permanently'), which clearly separates it from save_payment_link and list_payment_links. It does not explicitly name sibling tools or contrast against them, so it falls just short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by 'one of the user's saved payment links,' and the fallback sentence conveys a consequence of using it. However, there is no explicit when-to-use/when-not guidance and no named alternative (e.g., prefer save_payment_link over deleting).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_recurring_transactionRecurring Transactions: Delete recurring templateADestructiveInspect
Delete a recurring transaction template.
Already-spawned children are NOT touched (they're real
transactions in the user's history). Only the template stops
producing future instances.
Args:
template_id: The template Transaction.id
| Name | Required | Description | Default |
|---|---|---|---|
| template_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, but the description adds important scoping: only the template stops producing future instances, and existing child transactions remain untouched. This gives the agent a precise mental model of what will and will not be destroyed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: the core action appears first, the critical side-effect nuance follows immediately, and the parameter documentation is minimal and direct. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter deletion tool with a destructive annotation, a clear description of the target, and no output schema, nothing essential is missing. The description fully covers what the tool does, what it does not do, and how to supply the required identifier.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only an integer type with no description, so schema coverage is 0%. The description compensates by defining template_id as the template Transaction.id, clarifying that it is the recurring template's identifier rather than a child transaction ID or arbitrary ID.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the exact operation and resource: deleting a recurring transaction template. It further distinguishes this from deleting regular transactions by explicitly targeting the template rather than spawned children, which differentiates it from sibling tools like delete_transaction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when this tool is appropriate: when the goal is to stop a recurring template from producing future instances. It explicitly warns that already-spawned children are not touched, signaling that this is not the tool for removing individual child transactions, though it does not name the alternative tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_reminderAgent: Delete scheduled reminderBDestructiveInspect
Delete one of the user's scheduled reminders permanently.
Args:
reminder_id: The reminder's id.
Returns:
{"success": True, "deleted_id": ...} or an error when not found.
| Name | Required | Description | Default |
|---|---|---|---|
| reminder_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=false, and readOnlyHint=false, so the safety profile is covered. The description usefully adds that the deletion is permanent and that a not-found case returns an error, but says nothing about permissions or whether dependencies (e.g. notification state) are affected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Short and front-loaded, with the core action stated first and the argument/return details following in a docstring format. Slightly rigid Args/Returns scaffolding for a one-parameter tool, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description helpfully documents the return shape ({"success": True, "deleted_id": ...}) and the failure case. For a simple single-parameter delete whose destructive nature is already in annotations, the essentials are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description only says "reminder_id: The reminder's id," which essentially restates the parameter name. It does not clarify the integer type, where the id comes from (e.g. list_reminders), or how to recover from a bad id beyond noting an error is returned.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Delete one of the user's scheduled reminders") plus a manner qualifier ("permanently"). It does not, however, name or distinguish itself from relevant siblings such as create_reminder or list_reminders, so the agent gets a clear action but no routing signal.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use, when-not-to-use, or alternative guidance. Nothing tells the agent how this differs from list_reminders or whether a soft-delete alternative exists. The description assumes the caller already knows it wants to delete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_scenarioScenarios: Delete a what-if scenarioAInspect
Delete (archive) a what-if scenario.
Soft-delete — the scenario is hidden but retained.
Args:
scenario_id: The scenario to delete.
Returns:
{"success": True, "deleted_scenario_id": id} or {"error": "..."}.
| Name | Required | Description | Default |
|---|---|---|---|
| scenario_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without the description the agent would only see destructiveHint=false and idempotentHint=false; the text explains that this is a soft-delete that hides but retains the record, which resolves the otherwise confusing non-destructive annotation, and it also discloses the return payload. It stops short of saying whether archiving is reversible or who may perform it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded one-line purpose followed by the soft-delete caveat, then Args/Returns; every line earns its place, though the labeled Args/Returns scaffolding is slightly heavier than needed for a single scalar parameter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the inline Returns block usefully describes the success and error shapes, and the soft-delete note covers the main behavioral question for a one-parameter mutation. Missing only permission/authorization context and whether an archived scenario can be restored.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden, yet 'scenario_id: The scenario to delete' essentially restates the field name and adds no format, source (e.g. from list_scenarios), or validation detail. Minimal compensation for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (delete) plus the resource (what-if scenario) and immediately disambiguates by glossing the operation as 'archive', which separates it from hard deletes such as delete_transaction and from sibling scenario tools (get_scenario, update_scenario, commit_scenario).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the name; there is no statement of when to archive a scenario versus committing, updating, or simply listing it, and no prerequisites. Adequate but with a clear gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_tagTags: Delete tagADestructiveInspect
Delete a tag. The tag is removed from every transaction it was attached to, but the transactions themselves are untouched.
| Name | Required | Description | Default |
|---|---|---|---|
| tag_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, and the description adds meaningful context by disclosing that the tag is removed from every attached transaction while transactions themselves remain untouched. This goes beyond the annotation by clarifying the scope of the destructive behavior, which is valuable for an agent assessing impact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two tight sentences with zero filler. The primary action is front-loaded in the first sentence, and the second sentence adds an essential behavioral caveat without redundancy. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter destructive tool with annotations covering the safety profile, the description covers the key behavioral nuance (transaction preservation) and operation scope. It lacks explicit irreversibility notes, but destructiveHint already communicates that, and no output schema exists to be explained. The only minor gap is not stating error behavior for a non-existent tag.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of parameter explanation, but it does not explicitly describe tag_id. The parameter is a single obvious integer ID and the action 'delete a tag' strongly implies its meaning, so the gap is minor. Still, the description adds no direct parameter-level semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Delete a tag.' It further distinguishes itself from sibling delete tools by naming the exact resource (tag) and specifying the effect. The second sentence clarifies the scope, making it unambiguous among many delete_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is used to permanently remove a tag and explains the side effect on transactions, but it does not explicitly state when-not to use it or mention alternatives such as update_tag or list_tags. Usage context is clear from the name and effect, but no explicit routing guidance is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
delete_transactionTransactions: Delete transactionADestructiveInspect
Delete a transaction and reverse its balance effect on the account.
Args:
transaction_id: The ID of the transaction to delete
Returns:
Confirmation with updated account balance
| Name | Required | Description | Default |
|---|---|---|---|
| transaction_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false. The description adds value beyond these by specifying that the transaction's balance effect is reversed and that the return value is a confirmation with the updated account balance. This gives the agent a clear model of the operation's side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: the action is stated first, followed by an Args section and a Returns section. Every sentence contributes necessary information without repeating the title or restating the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive operation, the description covers the action, the balance effect, and the return format. Annotations already convey destructiveness, so no further behavioral disclosure is required. The absence of an output schema is compensated by the explicit Returns statement.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description includes an Args section defining transaction_id as 'The ID of the transaction to delete'. This provides the meaning the schema lacks, turning a bare integer into an actionable identifier. It is concise but sufficiently specific.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: delete a transaction. The additional clause 'reverse its balance effect on the account' clarifies the scope and consequence of deletion. This clearly distinguishes it from sibling tools like bulk_delete_transactions or undo_transaction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives such as bulk_delete_transactions or undo_transaction. It does not mention prerequisites, exclusions, or scenarios where another tool would be more appropriate. The usage context is left entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dismiss_budget_recommendationBudgets: Dismiss a budget recommendationAInspect
Snooze the spending-recommendation nudge for a budget for 30 days.
Mirrors the web "dismiss recommendation" action — sets the budget's
``recommendation_dismissed_until`` to 30 days from today so the
over/under-spending suggestion stops surfacing until then. Does not
change the budget amount or any spending data.
Args:
budget_id: The budget to snooze the recommendation for (the
caller's own, active).
Returns:
``{"success": True, "budget_id": ..., "dismissed_until":
"YYYY-MM-DD"}`` or ``{"error": "..."}`` if not found.
| Name | Required | Description | Default |
|---|---|---|---|
| budget_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=false). The description adds valuable specifics beyond that: it sets recommendation_dismissed_until to 30 days from today and explicitly does not change the budget amount or spending data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence front-loads the core action, and the remaining sentences add behavioral detail and return format without waste. The Args/Returns structure is appropriate because there is no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter mutation with no output schema, the description covers the action, exact field changed, duration, side-effect boundaries, parameter constraint, and return shape. Nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does by specifying that budget_id refers to the caller's own, active budget, adding a meaningful ownership/state constraint beyond the bare integer type in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (snooze/dismiss), resource (spending-recommendation nudge for a budget), and scope (30 days). It also maps to the exact field being set, making it easy to distinguish from other dismiss_* and update_budget siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what the tool does but offers no guidance on when to use it versus alternatives such as dismiss_payment_warning, dismiss_celebrations, or update_budget. There are no exclusions or prerequisites stated beyond the implicit 'caller's own, active' note.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dismiss_celebrationsGamification: Dismiss celebrationsAInspect
Mark financial milestones as celebrated (user has seen them).
Args:
milestone_ids: List of milestone IDs to dismiss
Returns:
Number of milestones marked as celebrated
| Name | Required | Description | Default |
|---|---|---|---|
| milestone_ids | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=false, destructiveHint=false, idempotentHint=false, so the agent knows this is a non-read, non-destructive, non-idempotent operation. The description adds that it marks milestones as celebrated and returns the count, which is useful. It doesn't disclose side effects like whether dismissed milestones can be re-dismissed or if this affects gamification state beyond visibility, but the annotations cover the basic safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: a one-sentence purpose, then Args and Returns sections. Every sentence earns its place, and the format is easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter tool with no output schema, the description covers the main action and return value. It lacks details about error behavior, whether the operation is idempotent (annotations say it isn't), and how it relates to the gamification system's celebration lifecycle. Given the simplicity, it's adequate but not rich.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does explain that milestone_ids is a list of milestone IDs to dismiss, which adds meaning beyond the raw schema. However, it doesn't specify constraints like whether IDs must be valid existing milestones, whether duplicates are allowed, or what happens if an ID is invalid. The description provides the core semantic but not edge-case guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Mark') and resource ('financial milestones as celebrated'), and clarifies the user-facing meaning ('user has seen them'). It is distinguishable from siblings like dismiss_budget_recommendation and dismiss_payment_warning by the resource type, though it doesn't explicitly name those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: call this when the user has seen a financial milestone and it should no longer be shown as a celebration. It does not explicitly state when not to use it or mention alternatives like get_celebrations or list_milestones, but the context is reasonably clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dismiss_payment_warningPlaid: Dismiss payment warningAInspect
Dismiss a stale "past due" payment warning on a debt account.
Applies when the dashboard shows a payment as past due but the user
knows the underlying Plaid data is wrong (e.g. they paid offline
and Plaid hasn't picked it up). Clears next_payment_due_date and
is_overdue; both repopulate on the next Plaid liabilities sync.
Args:
account_id: The Account.id to clear the warning on
Returns:
Status with the cleared fields.
| Name | Required | Description | Default |
|---|---|---|---|
| account_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing exactly what is mutated (clears next_payment_due_date and is_overdue) and that the change is transient because both repopulate on the next Plaid liabilities sync. That reversibility/regeneration detail is precisely the kind of behavioral context annotations don't supply, and nothing contradicts the hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the one-line purpose, then conditions, effects, args, and returns. Every sentence earns its place, though the blank-line-separated Args/Returns scaffolding is slightly more verbose than needed for a single-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a narrow mutation tool with annotations covering safety and no output schema, the description is complete: it explains the trigger, the exact fields cleared, the regeneration behavior, the sole argument, and the return shape. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single parameter (account_id) is undocumented in the schema, but the description compensates by explaining it is the Account.id whose warning should be cleared. Good compensation; minor room for the expected id format.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (dismiss) plus resource (past due payment warning) and scope (on a debt account). It is clearly distinguishable from siblings like confirm_debt_payment_match, reject_debt_payment_match, and dismiss_budget_recommendation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete triggering condition: dashboard shows past due but the user knows the Plaid data is wrong (e.g. paid offline, not yet picked up). This is strong when-to-use guidance, though it names no explicit alternative tool or exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
edit_family_chat_messageFamily: Edit family chat messageAInspect
Edit one of the user's own family chat messages. Other members'
messages can't be edited.
Args:
message_id: ID of the message to edit
content: New message text (1-500 chars)
Returns:
Updated message on success, or error dict
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | ||
| message_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only, non-idempotent, non-destructive mutation. The description adds meaningful context beyond that: an ownership/authorization constraint (only the user's own messages) and a content length limit (1-500 chars), plus a return contract. It doesn't cover reversibility or rate limits, but the added constraints are substantive given the annotation baseline.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with purpose and the ownership rule, then uses clean Args/Returns sections. It is appropriately sized with little waste, though the explicit Returns section is mildly redundant given the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter edit tool with no output schema, the description covers the purpose, the key authorization constraint, both parameter meanings, and the return value ("Updated message on success, or error dict"). Nothing an agent needs to invoke this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It compensates partially for content by adding a 1-500 character constraint not present in the schema, but message_id is only restated ("ID of the message to edit") with no format or source guidance. This is minimum viable rather than rich parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ("Edit one of the user's own family chat messages") and immediately distinguishes the operation from siblings like delete_family_chat_message and send_family_chat_message by stating ownership scope. An agent can tell exactly what this tool does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear usage condition and exclusion ("Other members' messages can't be edited"), which directly tells the agent when this tool is and isn't applicable. However, it does not name explicit alternatives (e.g., send_family_chat_message for new content), so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_accounts_csvAccounts: Export accounts CSVARead-onlyIdempotentInspect
Export all active accounts as CSV text.
Columns: Name, Account Type, Category, Balance, Currency, Institution,
Interest Rate, Credit Limit, Minimum Payment, Term (Months). Mirrors the
web "Export accounts" download and the REST data/export_accounts endpoint
(transactions have a separate CSV export).
Returns:
``{"csv_data": "<csv string incl. header row>", "row_count": N}`` —
active accounts only.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds real behavioral context beyond that: it exports only active accounts, enumerates the exact columns produced, and notes it mirrors the web download and REST data/export_accounts endpoint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then the column list and return contract. The column enumeration is slightly long but earns its place by telling the agent exactly what the CSV contains; no filler prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description supplies the return shape ({csv_data, row_count}) and clarifies the header row, so an agent knows what it will get back. Complete for a zero-parameter export tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. The column list is output semantics rather than input semantics, but it is still useful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (export) + resource (active accounts) + format (CSV text), and it explicitly distinguishes itself from the transactions CSV export sibling. An agent can tell this apart from export_transactions_csv and export_all_data without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the scope ('active accounts only') and points to the transactions CSV export as the separate path for transactions, which routes the agent away from a plausible mistake. It stops short of naming export_all_data as the alternative for a full dump, so it is clear context without full exclusion coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_all_dataReports: Export all data (accounts + transactions)ARead-onlyIdempotentInspect
Export the user's full data set — all accounts + all transactions —
as CSV strings plus a base64-encoded ZIP bundle. Premium-only.
Parity with the REST ``/api/v1/data/export_all/`` ZIP download. The
per-resource account and transaction CSV exports are free; this
Premium "export everything" call bundles both in one response.
``zip_base64`` decodes to a real .zip file; the CSV strings are
readable directly.
Returns:
``{"success": True, "accounts_csv": "...", "transactions_csv":
"...", "zip_base64": "...", "filename": "zoninga_export_YYYY-MM-DD.zip",
"account_count": N, "transaction_count": M}``, or
``{"error": "..."}`` if not Premium.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, but the description adds behavior beyond them: the Premium-only gate, the error shape returned when the user is not Premium, and the fact that zip_base64 decodes to a real .zip while the CSV strings are directly readable. That is real, non-redundant context for a zero-parameter export.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and scope in the first sentence, followed by the Premium/parity caveat and a structured Returns block. Slightly verbose with backtick formatting and the REST-parity sentence, but every sentence carries useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description carries the return contract itself and does so fully: success payload fields, the filename pattern, counts, and the error case. An agent has everything needed to call and interpret the result.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the baseline is 4; there is nothing for the schema to document and nothing the description needs to compensate for. No parameter-related guidance is missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with explicit scope: the user's full data set (all accounts + all transactions) delivered as CSV strings plus a base64 ZIP. It also names the exact sibling alternatives (export_accounts_csv / export_transactions_csv) and explains how this tool differs from them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells the agent when to pick this over the free per-resource exports ('The per-resource account and transaction CSV exports are free; this Premium "export everything" call bundles both in one response') and discloses the Premium precondition. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
export_transactions_csvTransactions: Export transactions as CSVARead-onlyIdempotentInspect
Export transactions as CSV data.
Args:
account_id: Filter to a specific account (optional)
start_date: Start date YYYY-MM-DD (optional)
end_date: End date YYYY-MM-DD (optional)
Returns:
CSV data as a string with column headers
| Name | Required | Description | Default |
|---|---|---|---|
| end_date | No | ||
| account_id | No | ||
| start_date | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, covering the safety profile. The description adds the useful behavior that the tool returns 'CSV data as a string with column headers', which is a non-obvious output trait. It does not discuss edge cases like date inclusivity or row limits, but the annotation coverage lowers the burden.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well organized with clear sections: one-sentence purpose, parameter list, and return statement. No redundant or filler content is present, and each line earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity, read-only, optional-parameter export tool, the description provides sufficient information: the action, the filters, and the return type. It does not list exact CSV headers or when to prefer sibling export tools, but these are minor gaps given the simple surface area and helpful annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries full responsibility for explaining parameters. It provides a one-line meaning for each parameter and specifies the date format (YYYY-MM-DD). It could clarify default behavior when no filters are given, but it already adds meaningful semantics beyond the bare schema properties.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Export'), resource ('transactions'), and output form ('CSV data'), which clearly differentiates it from siblings like export_accounts_csv and export_all_data. The title reinforces the same scope. An agent can identify what this tool does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The optional filter parameters imply usage for exporting transactions by account and date range, but the description does not explicitly say when to choose this over list_transactions, export_accounts_csv, or export_all_data. There are no stated exclusions or alternative-routing guidance, so usage is mostly implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
forecast_cash_flowAnalytics: Forecast cash flowARead-onlyIdempotentInspect
Project cash flow for the next N months based on recent patterns.
Uses 3-month averages of income and expenses as a baseline AND
biases each future month by recurring templates the user has
scheduled but that haven't fired yet (Phase 115). Templates that
have already fired at least once are folded into the historical
baseline, so they're not counted twice.
Each projection row carries ``scheduled_income_added`` and
``scheduled_expenses_added``, which separate the part of the
projection that comes from "what's already happening" from
"what's scheduled to start." The top-level
``scheduled_templates_count`` is the number of un-fired templates
contributing to the bias — when 0, the projection is purely
historical.
Args:
months_ahead: Number of months to project (default 3, max 12).
Values <= 0 return a validation error; values > 12 are clamped
to 12 and the response sets clamped_months_ahead=true while
requested_months_ahead echoes the original input.
account_id: Optional - project for a specific account only
Returns:
Current balance, average monthly income/expenses/net,
monthly projections (each with the scheduled-bias breakout),
``scheduled_templates_count``, and a human-readable note.
| Name | Required | Description | Default |
|---|---|---|---|
| account_id | No | ||
| months_ahead | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive behavior, so the burden is lower. The description goes beyond them by disclosing the double-counting rule for already-fired templates, the scheduled_income_added/expenses_added breakout, and the clamped_months_ahead / validation-error edge cases — genuinely useful behavior beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then structured Args/Returns sections. It is verbose — the Phase 115 reference and extended prose on the bias mechanism are heavier than strictly necessary — but each part informs correct invocation of an analytical tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully documents the return shape (balance, averages, projections, scheduled_templates_count, note) and covers both parameters. It is nearly complete for this analytical tool; only corner cases like empty history/no accounts are unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries full load. months_ahead is richly specified (default 3, max 12, <=0 error, >12 clamped with a response flag), which compensates well; account_id is thinner ('optional - project for a specific account only') but adequate given the integer type is in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Project cash flow for the next N months') and immediately scopes it as forward-looking, which implicitly separates it from the historical sibling get_cash_flow. An agent can tell what this does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the forecasting methodology in detail, which implies when the tool is relevant, but it never explicitly states when to pick this over alternatives like get_cash_flow or get_period_summary, nor any exclusions. Usage is left to inference from the purpose.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_account_detailsAccounts: Get account detailsBRead-onlyIdempotentInspect
Get detailed information about a specific account.
Args:
account_id: The ID of the account to retrieve
Returns:
Account details including recent transactions.
Debt accounts include Plaid enrichment: is_overdue, loan_status,
has_pmi, has_prepayment_penalty, last_statement_balance,
next_payment_due_date, last_payment_amount/date,
ytd_interest/principal_paid, and credit card APR breakdown.
| Name | Required | Description | Default |
|---|---|---|---|
| account_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, covering the safety profile. The description adds behavioral value by specifying the return shape: recent transactions and conditional Plaid enrichment for debt accounts, which is not present in annotations or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with Args and Returns sections and contains no filler. The Returns list is long but justified by the many debt-account enrichment fields. The core purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only tool with no output schema, it covers the main return values and debt-account specifics. However, it lacks error behavior and account_id provenance, and could state whether recent transactions are limited in count or date range.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and the description's Args section only says 'The ID of the account to retrieve,' which adds little beyond the parameter name and type. It does not explain where the ID comes from, validity constraints, or behavior for missing accounts.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Get detailed information about a specific account.' This clearly distinguishes from list_accounts by requiring a specific account_id. However, it does not explicitly differentiate from other get_* siblings like get_account_health or get_account_history.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as list_accounts, get_account_health, or get_account_history. The description implies usage through the verb, but offers no prerequisites, exclusions, or alternate routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_account_healthAnalytics: Get account healthARead-onlyIdempotentInspect
Health check across all accounts.
Identifies accounts that need attention: high credit utilization,
no recent activity, approaching limits.
Returns:
List of accounts with status (healthy/warning/critical),
balance, credit utilization, interest rate, days since
last transaction.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the description doesn't need to repeat safety behavior. It adds valuable behavioral context by specifying the output structure (list of accounts with status, balance, credit utilization, interest rate, days since last transaction) and the criteria it evaluates. This is beyond what annotations and schema (empty) provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-organized: a one-line purpose, a sentence listing attention criteria, and a 'Returns' section enumerating output fields. Every sentence earns its place, and the most important information is front-loaded. It is neither padded nor under-specified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only tool with no output schema, the description is quite complete. It specifies what the tool does, what it returns, and the criteria used. Minor omissions like whether all account types (e.g., closed accounts) are included or whether results are paginated are not critical given the tool's apparent simplicity and the absence of parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema coverage is 100% (empty schema). The description adds the scope 'across all accounts', which is the only meaningful semantic. Baseline for zero parameters is 4; the description appropriately clarifies that no filtering is available.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action: 'Health check across all accounts' and then details what it does: identifies accounts needing attention with concrete criteria (high credit utilization, no recent activity, approaching limits). It distinguishes itself from siblings like get_connection_health or get_financial_health_score by focusing on per-account status rather than connection health or an aggregate score.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use this tool: when you need a health overview of all accounts and want to locate problem accounts. However, it does not explicitly mention alternatives or exclusions, such as using get_account_details for a single account or get_financial_health_score for a single score. The usage is implied but not explicitly contrasted with siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_account_historyAccounts: Get account historyARead-onlyIdempotentInspect
Get transaction history for a specific account with running balance.
Shows transactions in reverse chronological order with a running
balance column, useful for seeing how the balance changed over time.
By default, only transactions dated **today or earlier** are returned
— this keeps the running-balance trail anchored to the live
``current_balance``. Pass ``include_scheduled=True`` to also see
future-dated rows (e.g. recurring transactions the cron will spawn);
when scheduled rows are included, the running-balance walk anchors at
the projected balance (``current_balance`` plus the sum of unapplied
future impacts) so the math still reconciles, and each future row
carries an ``is_scheduled: True`` flag.
Args:
account_id: Account to show history for
limit: Maximum transactions to return (default 50, max 200)
include_scheduled: When False (default), filter out future-dated
rows. When True, include them and walk from the projected
balance.
Returns:
Account info and transaction history with running balances.
Response includes ``include_scheduled`` so callers know which
view they got and ``projected_balance`` when scheduled rows
are included.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| account_id | Yes | ||
| include_scheduled | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description goes beyond this by explaining default date filtering, how include_scheduled changes the running-balance anchor to the projected balance, and the is_scheduled flag on future rows. It also tells callers the response will echo include_scheduled and include projected_balance, providing concrete behavioral detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The main purpose is front-loaded in the first sentence, with default behavior, parameter details, and return values following in logical order. Every sentence adds value, and the longer include_scheduled explanation is justified because it covers subtle running-balance math.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the Returns section adequately describes what callers get: account info, transaction history with running balances, the echoed include_scheduled flag, and projected_balance when scheduled rows are included. The description also covers defaults and edge behavior, making it complete enough for an agent to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden, and it fully compensates. The Args section explains account_id, limit (including the max of 200 that the schema omits), and include_scheduled in behavioral terms. This is more informative than the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Get transaction history for a specific account with running balance.' It also explains the return orientation and running-balance column, making the tool's purpose unambiguous. The name and description together clearly distinguish this from shared-account or plain list-transaction tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: viewing an individual account's transaction history with a running balance, and it explains the default date filter and the include_scheduled option. It does not explicitly name alternatives like get_shared_account_history or get_account_details, so it stops short of full when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_agent_chat_sessionAgent Chat: Get agent chat sessionARead-onlyIdempotentInspect
Get the full transcript of a past AI assistant chat session.
Returns the messages the user exchanged with the in-app Zoninga
Assistant. Useful for summarizing a past conversation or
continuing a thread from a prior session.
Args:
session_id: The chat session's ID.
Returns:
{"session": {...}, "messages": [{role, content, created_at}, ...]}
or {"error": "..."} if not owned by the caller.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a full safety profile (readOnly, idempotent, non-destructive, closed-world), so the bar is low. The description goes beyond them by disclosing the ownership constraint – returning {"error": "..."} if the session is not owned by the caller – which an agent cannot learn from annotations and needs to handle failures.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences cover purpose, use cases, and return shape, followed by compact Args/Returns blocks. Every line carries information and the most important statement (what the tool returns) comes first; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by sketching the return payload ({"session": ..., "messages": [{role, content, created_at}]}) and the error case, and the single parameter is acknowledged. The only meaningful omission is how to obtain a valid session_id, which the agent must infer from the sibling list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the schema only supplies a title, so the description's 'session_id: The chat session's ID' is nearly tautological with the parameter name. It does not state the type (integer), the format, or where a valid ID comes from (list_agent_chat_sessions), leaving the parameter thinly documented on both sides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Get the full transcript of a past AI assistant chat session', and the second sentence pins down exactly what is returned (messages exchanged with the in-app assistant). This implicitly separates it from the list_agent_chat_sessions sibling, but it never names that sibling or any other tool, so an agent must infer the boundary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Useful for summarizing a past conversation or continuing a thread from a prior session' gives concrete usage contexts, which is better than most siblings. It stops short of a 5 because there are no explicit alternatives named (e.g., use list_agent_chat_sessions first to obtain a session_id) and no when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_agent_instructionsAgent: Get agent operating instructionsARead-onlyIdempotentInspect
Get behavioral guidelines, safety rules, and workflow guidance for the AI agent.
Describes how the assistant operates as the user's personal finance
assistant; intended for the start of a conversation.
Returns:
Structured instructions including role, safety rules, workflows, and limits
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive, closed-world behavior, so the safety profile is covered. The description adds genuine value by describing the returned content (role, safety rules, workflows, limits), which matters because there is no output schema. It does not mention auth or rate-limit behavior, keeping it short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose followed by a use-case note and a return summary; each line earns its place, especially the return breakdown given the missing output schema. The indented 'Returns:' block is slightly verbose but not wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only tool with rich annotations and no output schema, the definition is nearly complete: it explains what the call yields. Only minor gaps remain, such as whether the content is static or personalized to the user.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is nothing for the description to disambiguate, and it correctly says nothing about inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get behavioral guidelines, safety rules, and workflow guidance for the AI agent') and elaborates that it describes how the assistant operates as a personal finance assistant. This distinguishes it from lookalikes such as get_user_preferences or get_user_profile, though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit usage moment: 'intended for the start of a conversation.' There is no when-not guidance or named alternative, but for a zero-parameter bootstrap tool that is largely unnecessary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_balance_sheetReports: Get balance sheetBRead-onlyIdempotentInspect
Get a balance sheet snapshot showing assets, liabilities, and net worth.
Returns asset and liability breakdowns by account with totals.
Returns:
Balance sheet data with asset/liability detail and net worth.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so safety is covered. The description adds that the result is a 'snapshot' and includes breakdowns by account with totals, which is useful but does not address details like as-of date or account scope; still, it does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first two sentences are efficient, but the third sentence ('Returns: Balance sheet data with asset/liability detail and net worth.') merely repeats information already present in the first and second sentences. This redundant closing line wastes space and lowers overall conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only report, the description covers what is returned and the key components (assets, liabilities, net worth). It does not, however, explain how this differs from the similarly named get_net_worth tool, and the absence of an output schema places more burden on the description to define the result structure, which is only partially met.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema provides complete coverage. The description adds no parameter-specific semantics, but none are needed; baseline 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description opens with 'Get a balance sheet snapshot showing assets, liabilities, and net worth,' a specific verb and resource that clearly identifies the financial report. It does not explicitly name sibling alternatives like get_net_worth or get_income_statement, but the balance sheet concept is distinct and agents familiar with accounting terms will select it correctly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus related reporting tools such as get_net_worth or get_income_statement. No exclusions or alternative tool mentions appear, leaving the agent to infer the distinction from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_budget_historyAnalytics: Get budget historyBRead-onlyIdempotentInspect
Get budget performance over time (actual vs budget per month).
Shows utilization percentage, over/under status for each month.
Args:
budget_id: Budget ID (provide this OR category_name)
category_name: Category name of the budget (provide this OR budget_id)
months: Number of months (default 6)
Returns:
Monthly budget amount, actual spent, utilization %, over/under status.
| Name | Required | Description | Default |
|---|---|---|---|
| months | No | ||
| budget_id | No | ||
| category_name | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so safety is covered. The description adds the default months value and the output fields, but not deeper behavior like what happens when both identifiers are provided or the history limit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well organized with purpose, output summary, args, and returns. Slight redundancy between 'Shows utilization percentage, over/under status' and the Returns section, but otherwise no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Adequate for a read-only query tool: it documents all parameters and return values. Missing edge cases like providing both identifiers, neither identifier, or invalid month counts, but an agent can likely call it correctly with the given information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the full burden. It clearly explains each parameter, especially the OR relationship between budget_id and category_name, and the default value of months, adding meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the verb 'get', the resource 'budget performance', and the scope 'over time (actual vs budget per month)', which clearly distinguishes it from budget streaks or category trend tools. It doesn't explicitly name sibling tools, but the budget-vs-actual focus is specific enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus other analytics tools such as get_budget_streaks or get_category_trend. The 'provide this OR category_name' instruction is parameter usage, not tool-selection context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_budget_streaksGamification: Get budget streaksARead-onlyIdempotentInspect
Get budget streak data: consecutive months on-budget per category.
Returns:
List of budget streaks with category, consecutive months, and longest streak
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering safety and side-effect behavior. The description adds the return format (list of streaks with fields) but does not disclose other behavioral traits such as pagination, sorting, or potential data freshness. Given the annotations carry the safety burden, the description's additional context is adequate but not rich, so a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is exceptionally concise: two lines plus a returns block. It front-loads the core purpose, then specifies the output structure. Every sentence earns its place, and there is zero waste or repetition. The structure is clean and immediately scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with no parameters, no output schema, and safety annotations already in place, the description covers all necessary aspects: what it does, what it returns, and the definition of a streak. There is no missing information an agent would need to invoke it correctly. It is fully complete for its complexity level.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description has nothing to add regarding parameter meaning. The rubric assigns a baseline of 4 for 0-parameter tools, and the description does not introduce any confusion or missing information. It correctly remains silent on parameters since there are none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Get budget streak data: consecutive months on-budget per category.' It names a specific resource (budget streaks), defines the concept, and specifies granularity (per category). It also lists the return fields, making the tool's function unmistakable. Among many get_* siblings, this one is uniquely identified by its focus on streaks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention exclusions, conditions, or how it differs from related gamification tools like get_gamification_summary, get_points_and_level, or get_leaderboard. An agent would have to infer its usage from the name and description alone, which is insufficient for optimal tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_cash_flowReports: Get cash-flow reportARead-onlyIdempotentInspect
Get cash flow analysis (income vs expenses).
Args:
start_date: Start date in YYYY-MM-DD format (defaults to 30 days ago)
end_date: End date in YYYY-MM-DD format (defaults to today)
include_investment_deposits: When False (default), 401(k)/IRA/
brokerage deposits are excluded from income — disposable
cash flow view. When True, they're counted as income —
useful for "total wealth flow" or "where did all my money
go" reporting that includes retirement contributions.
Returns:
Income and expense totals with net cash flow. Includes
``include_investment_deposits`` so the caller knows which view
it just received.
| Name | Required | Description | Default |
|---|---|---|---|
| end_date | No | ||
| start_date | No | ||
| include_investment_deposits | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive, so the bar is lower. The description adds useful transparency by explaining the two views and that the response echoes include_investment_deposits so the caller knows which view was generated, which helps set expectations for returned data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description opens with a one-line purpose, then systematically covers Args and Returns. Every sentence adds value, and the structure mirrors the schema, making it easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only report with no output schema, the description covers parameters, return contents, and a key behavioral flag. It doesn't explicitly address how this differs from similar reporting siblings, but the core invocation context is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description fully compensates: it specifies YYYY-MM-DD format and defaults for both dates aligned with schema defaults, and gives a detailed, nuanced explanation of include_investment_deposits with concrete use cases for each value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states 'Get cash flow analysis (income vs expenses),' naming a specific verb and resource. However, it does not explicitly distinguish itself from sibling reports like get_income_statement or forecast_cash_flow, leaving some inference to the agent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides strong guidance on the include_investment_deposits parameter (disposable vs total wealth view) but gives no explicit direction on when to select this tool over sibling report tools. Usage context is implied by the title and description rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_category_suggestionsCategories: Get category suggestionsARead-onlyIdempotentInspect
Get smart category suggestions for a transaction description.
Uses a 3-strategy engine: user history, Plaid mapping, and keyword rules.
Useful when creating transactions to auto-suggest the right category.
Args:
description: Transaction description to match (e.g., "Starbucks", "Amazon")
transaction_type: Optional filter — "income" or "expense"
Returns:
Ranked list of suggested categories
| Name | Required | Description | Default |
|---|---|---|---|
| description | Yes | ||
| transaction_type | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is clear. The description adds valuable behavioral context by revealing the 3-strategy engine (user history, Plaid mapping, keyword rules), which helps the agent understand how suggestions are generated and why results may vary. It also notes the ranked list return, which is useful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: a one-sentence summary, a brief note on the engine, a usage context line, and clear Args/Returns sections. Every sentence earns its place, and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only suggestion tool with 2 simple parameters and no output schema, the description covers the essentials: what it does, how it works, when to use it, and what it returns. It could mention that the ranked list is limited in size or that suggestions are based on the user's own data, but these are minor gaps given the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does: it explains that 'description' is the transaction description to match and gives examples ('Starbucks', 'Amazon'), and it clarifies that 'transaction_type' is an optional filter with 'income' or 'expense' values. This adds meaning beyond the bare schema, though it doesn't detail the exact format of the ranked list.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Get smart category suggestions for a transaction description.' It specifies the resource (category suggestions) and the action (get), and distinguishes it from related tools like list_categories and bulk_categorize_transactions by focusing on suggestion generation for a single description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description says 'Useful when creating transactions to auto-suggest the right category,' which provides clear context for when to use it. It doesn't explicitly mention alternatives or when not to use it, but the use case is specific enough that an agent can infer when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_category_trendAnalytics: Get category trendARead-onlyIdempotentInspect
Get monthly spending trend for a specific category.
Shows how spending in a category has changed over time,
including percentage of total spending and month-over-month change.
Args:
category_id: Category ID (provide this OR category_name)
category_name: Category name, case-insensitive (provide this OR category_id)
months: Number of months (default 6)
Returns:
Monthly amounts, percent of total spending, and change percentages.
| Name | Required | Description | Default |
|---|---|---|---|
| months | No | ||
| category_id | No | ||
| category_name | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds behavioral context beyond annotations: it discloses exactly what the output contains (monthly amounts, percent of total spending, month-over-month change) and the mutual-exclusion constraint between category_id and category_name. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: one sentence states the core purpose, a second expands on what is shown, and a structured Args/Returns section follows. Every sentence earns its place with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema present, the description appropriately explains the return values (monthly amounts, percent of total, change percentages) and all parameters. The combination of purpose, parameter constraints, defaults, and return values makes the tool fully callable by an agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for documenting parameters. It explains each parameter individually, including the OR relationship between category_id and category_name, case-insensitivity, and the default for months. This is exactly the meaning the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description leads with a specific, specific verb+resource phrase: 'Get monthly spending trend for a specific category.' It clearly distinguishes this analytics tool from siblings like get_spending_summary or get_cash_flow by naming the exact analytical output (category trend). Full clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies its use case — retrieving monthly trend data for a specific category — which implies the context for choosing this tool among analytics siblings. It does not explicitly name alternatives or exclusions, but the context is unambiguous enough for an agent to select it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_celebrationsGamification: Get pending celebrationsARead-onlyIdempotentInspect
Get uncelebrated financial milestones (for celebration animations).
Returns:
List of milestone achievements the user hasn't seen yet
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly, idempotent, and non-destructive behavioreding. The description adds meaningful state-based context beyond those annotations by clarifying that only uncelebrated, unseen achievements are returned. It does not contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loads the core action before the return clarification. Every sentence earns its place, with no filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only tool, the description is largely complete: it states the input is none and the output is a list of unseen milestone achievements. It could be slightly stronger by noting whether fetching marks them as seen or whether dismissing must be handled separately, given the sibling tool dismiss_celebrations exists.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the schema has no parameter details to provide. The baseline of 4 applies here because the description accurately frames the tool's input-free invocation and reinforces that it returns a list without requiring arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get') and a specific resource ('uncelebrated financial milestones'), while the return clause clarifies these are milestone achievements the user hasn't seen yet. The phrase 'for celebration animations' adds context and helps distinguish this from plain milestone listing tools like list_milestones.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to call this tool: to obtain pending, uncelebrated milestones for celebration animations. It does not explicitly name alternatives or when-not-to-use conditions, but the pending/unseen scope is enough to separate it from related sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_connection_healthPlaid: Check connection healthARead-onlyIdempotentInspect
Per-connection staleness report: how long since each linked bank
last synced, and exactly which surfaces a stale connection taints.
Answers "is my data up to date?" and shows whether balance, budget
or forecast figures rest on stale bank data. A connection is
flagged stale after 14 days without a successful sync
(threshold_days in the response).
Returns:
``{"items": [...], "stale_count": N, "threshold_days": 14}``.
Each item: institution_name, last_synced_at, days_since_sync,
is_stale, error state, linked accounts (with per-account
days_since_balance_update), affected_budgets (budget names now
missing activity), affects_forecast, and reconnect_url — a
BROWSER deep link. The reconnect runs in the user's browser
(Plaid Link with step-up auth); this tool returns the link.
A closed account can be archived instead, and a manual sync
can refresh a connection that is healthy but merely behind.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnly/idempotent/non-destructive), the description discloses the 14-day staleness threshold, that the reconnect_url is a BROWSER deep link requiring Plaid Link step-up auth, and critically that this tool only returns the link rather than performing the reconnection. That boundary is exactly the kind of behavioral context an agent needs to avoid misusing the result.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose and usage framing are front-loaded in the first two paragraphs, and the Returns block is dense but earns its space because there is no output schema. The verbose field enumeration is the only mild bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully carries the burden of documenting the return payload — top-level keys, per-item fields (institution_name, last_synced_at, days_since_sync, is_stale, error state, per-account days_since_balance_update), affected_budgets, affects_forecast, and reconnect_url. Nothing an agent needs to interpret the result is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4; there is no parameter syntax for the description to clarify. It simply operates on all connections by default, which is consistent with the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb+resource combination — a per-connection staleness report covering time since each linked bank last synced and which surfaces (balance, budget, forecast) a stale connection taints. This is unmistakably distinct from siblings like get_plaid_item, list_plaid_items, or sync_plaid_item, even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It anchors usage with 'Answers "is my data up to date?"' and explains what the results reveal about downstream figures. It also points to remedies (reconnect via link, archive a closed account, manual sync for a healthy-but-behind connection), though it does not name the corresponding sibling tools or state when this tool should be skipped.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_daily_averagesAnalytics: Get daily averagesARead-onlyIdempotentInspect
Get daily average income, spending, and net for a date range.
Computes the averages server-side, handling partial months and
year-to-date ranges correctly, and provides a monthly breakdown.
Phase 114 — the start date is **clamped to the user's first
transaction date** when that's later than what was requested.
Without this, a user with two months of activity who asks for
the YTD average sees the daily averages roughly halved (divided
by ~120 days when only ~60 had activity). The response carries
``clamped_to_first_activity`` (bool) and ``requested_start_date``.
When ``clamped_to_first_activity`` is True, the averages cover
only the days since the first transaction (e.g. 55 days of data),
not the full requested range.
Args:
start_date: YYYY-MM-DD (defaults to Jan 1 of current year)
end_date: YYYY-MM-DD (defaults to today)
Returns:
Daily averages with monthly breakdown showing income, expenses,
and net for each month in the range, plus the two clamping
metadata fields.
| Name | Required | Description | Default |
|---|---|---|---|
| end_date | No | ||
| start_date | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), yet the description adds real value: server-side computation, partial-month/YTD handling, and the Phase 114 start-date clamping rule with its rationale and the returned metadata fields. It stops short of mentioning rounding, currency, or timezone behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose, then Args/Returns sections that are easy to scan. The clamping explanation is repeated (the halving example, then a restatement of the same effect), which is mild redundancy in an otherwise well-organized description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly describes the return payload (monthly income/expense/net breakdown plus the two clamping metadata fields), and params are documented. Completeness is good for the tool's complexity, with minor gaps around numeric formatting and currency assumptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the load, and it does: it documents both parameters with format (YYYY-MM-DD) and defaults (Jan 1 of current year; today). That fully compensates for the bare schema, though it adds no constraint/interaction detail between the two dates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource plus the computed metrics (daily average income, spending, net) and the scoping dimension (date range). It is clearly distinguishable from most siblings, though it never names an alternative analytics tool to contrast against.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when the tool applies (YTD/partial-month averaging, default Jan 1–today) but gives no explicit when-to-use/when-not guidance and never routes the agent to a sibling such as get_period_summary or get_cash_flow. Usage context is inferable but not stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_debt_paydownReports: Get debt-paydown analysisARead-onlyIdempotentInspect
Get a debt paydown analysis with payoff projections.
A call with no arguments returns the user's SAVED plan — their chosen
strategy plus their saved extra monthly payment, the same numbers the
Debt Paydown page shows them. Answers questions like "when will I be
debt-free?" / "what's my payoff date?": `debt_free_date` is the
report's computed payoff date, and `strategy` is the plan's strategy.
Passing `strategy` / `extra_monthly` produces a what-if projection
("what if I paid $200 extra?") rather than the saved plan.
Args:
extra_monthly: Extra monthly payment beyond minimums. Omit to use
the user's saved extra payment (their plan).
strategy: Payoff strategy - highest_rate, snowball, variable_snowball,
minimum_payments, highest_balance, highest_payment,
cashflow_index, npv, max_interest_savings. Omit to use
the user's saved plan strategy.
respect_goal_priorities: When True (default), active debt-payoff
goals override strategy ordering for the primary projection.
False shows what the chosen strategy does without goal
interference. The strategies_comparison table shows pure
strategy ordering regardless of this flag.
`this_month` answers "what do I pay this month?" / "which debt
gets my extra?": `this_month.target` is the debt the plan attacks
first (`target_basis`: extra_this_month, or first_rollover when no
extra is set and freed-up payments start rolling to it in
`first_extra_month`). Each row has payment (the amount to send =
minimum_payment + extra), minimum_payment (the account's stated
minimum), and plan_minimum_payment / plan_payment (the projection's
amortized figures, which can differ by a few dollars).
Returns:
Debt payoff timeline (incl. debt_free_date, the computed payoff date),
this_month (per-debt payments + the target debt), interest
savings, strategy comparison, and is_user_saved_plan (True when
the projection is the user's own saved plan).
| Name | Required | Description | Default |
|---|---|---|---|
| strategy | No | ||
| extra_monthly | No | ||
| respect_goal_priorities | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/destructiveHint=false, and the description adds substantial context beyond them: the saved-plan-vs-what-if distinction, respect_goal_priorities overriding strategy ordering, why plan_minimum_payment can differ from minimum_payment, and the is_user_saved_plan output flag. This is rich behavioral disclosure for a read-only report.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well front-loaded with the core saved-plan/what-if distinction, then organized into Args and Returns sections that each earn their place. It is dense and slightly long with a few embedded asides (e.g., the amortized-figure caveat), but there is little outright waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description steps in to name the key return fields (debt_free_date, strategy, this_month.target/target_basis, first_extra_month, is_user_saved_plan) and their meaning. Combined with the parameter and mode coverage, an agent has everything needed to call and interpret this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the full load, and it does: extra_monthly is defined as payment beyond minimums with omit-behavior, strategy enumerates nine valid values that are absent from the schema, and respect_goal_priorities explains the True-default semantics and the strategies_comparison exception. Nothing about the parameters is left undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get a debt paydown analysis with payoff projections') and immediately pins down the two operating modes: no-arg returns the saved plan, while passing strategy/extra_monthly produces a what-if projection. No sibling among the ~130 tools offers debt-payoff projection, so the scope is unambiguous without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context-based guidance: omit arguments for the saved plan, pass strategy/extra_monthly for a what-if, and it ties the call to concrete user questions ('when will I be debt-free?', 'what do I pay this month?'). It never names an alternative tool or an explicit when-not-to-use condition, so it stops short of the top mark.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_family_activityFamily: Get family activity feedBRead-onlyIdempotentInspect
Get recent family activity feed.
Args:
limit: Maximum entries to return (default: 20, max: 100)
Returns:
List of recent family activities
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior, so the safety profile is covered. The description adds minimal context by mentioning 'recent' and the limit parameter, but it does not disclose what constitutes an activity, the ordering, or whether a family group must exist. It does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and well structured with Args and Returns sections. Every line carries relevant information and there is no filler or repetition of the annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the single parameter and read-only annotations, the description is adequate for the basic call. However, with no output schema, 'List of recent family activities' leaves the shape and scope of each activity unspecified, and the ambiguous 'activity' term could make selection among the many family-related tools harder.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It documents limit with a default of 20 and a maximum of 100, which adds information beyond the schema (which only includes the default). For the single parameter, this is sufficient and useful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'Get recent family activity feed.' It is not a tautology and names the specific object of interest. However, it does not explicitly distinguish itself from sibling getters/listers like get_family_group or list_family_members, and 'activity' remains somewhat underspecified.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives. There is no mention of preferred use cases, exclusions, or a reference to sibling tools such as list_family_members or get_family_group. Usage must be inferred entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_family_groupFamily: Get family groupARead-onlyIdempotentInspect
Get the user's family group info, members, and role.
Returns:
Family group details with member list, or message if not in a group
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, and non-destructive behavior. The description adds value by disclosing the return format ('family group details with member list') and the edge case of not being in a group ('message if not in a group'), which are useful behavioral details beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no wasted words. It front-loads the purpose and then states the return behavior, making it efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple parameterless read tool, the description provides the essential return information and the not-in-group case. Without an output schema, it could be slightly more detailed about the exact fields in the member list, but it is sufficiently complete for an agent to understand the tool's purpose and output.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the description does not need to explain parameter semantics. The schema coverage is trivially 100%, and the description does not add anything beyond the schema, but the baseline of 4 for a parameterless tool is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('Get') and resource ('family group info, members, and role'), which distinguishes it from more specific siblings like list_family_members. It is concise and unambiguous, though it could be slightly more explicit about the exact scope of 'info'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for retrieving the user's family group details, but it does not explicitly differentiate it from related tools like list_family_members or get_family_activity. No exclusions or alternative routing are mentioned, so usage guidance is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_financial_health_scoreReports: Get financial health scoreARead-onlyIdempotentInspect
Calculate a financial health score based on various metrics.
Returns:
Financial health score (0-100) with component breakdowns
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds that it returns a financial health score from 0 to 100 with component breakdowns, which is useful beyond the annotations. It does not disclose any side effects, but none are expected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely compact: one sentence plus a Returns line. It is front-loaded with the purpose and contains no filler, with every sentence providing useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description gives the return range and mentions component breakdowns, but it is vague about what metrics are used and whether the score applies to all accounts or a specific scope. With no output schema or parameters, more detail about the components and data prerequisites would make it more complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and the schema is empty with 100% schema description coverage, so there is nothing to add. The baseline of 4 applies because no parameter documentation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as calculating a financial health score, which is a specific verb-resource combination. However, it does not differentiate this tool from siblings like get_account_health or get_financial_ratios, and 'based on various metrics' is vague about the exact inputs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives. The description does not mention any conditions, prerequisites, or exclusions, leaving an agent without context among a long list of get_* report tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_financial_ratiosAnalytics: Get financial ratiosARead-onlyIdempotentInspect
Get key financial ratios computed from the user's data.
Returns pre-computed ratios.
Three different "debt-to-income" flavors are returned and they are
not interchangeable; each fits a different use case:
Returns:
- savings_rate_pct: (income - expenses) / income * 100
- expense_ratio_pct: expenses / income * 100
- debt_to_income_ratio: BALANCE-SHEET LEVERAGE RATIO —
total liabilities / annual income, as a multiple (e.g. 2.36).
A normal homeowner sits above 1.0 because mortgage balance
dwarfs annual income; that's expected for this metric and
NOT a sign of financial trouble. Suited to net-worth analysis.
- back_end_dti_pct: MORTGAGE-INDUSTRY BACK-END DTI —
monthly debt service / monthly gross income, as a percent.
28% / 36% are the GSE qualifying-mortgage thresholds. Same
calc the health-score DTI component uses. None when no
debt accounts have ``minimum_payment`` populated.
- front_end_dti_pct: MORTGAGE-INDUSTRY FRONT-END DTI —
housing payment (mortgage P&I + escrow tax + escrow
insurance) / monthly gross income, as a percent. 28% is
the standard threshold. None when the user has no mortgage
accounts. Measures housing affordability, as opposed to
the "all debt" back-end view.
- monthly_debt_service: dollar sum of minimum payments across
active debt accounts (numerator of back_end_dti_pct).
- monthly_housing_payment: dollar sum of mortgage P&I + escrow
across active mortgage accounts (numerator of front_end_dti_pct).
- emergency_fund_months: liquid assets / monthly expenses
- credit_utilization_pct: CC balances / CC limits * 100
- financial_assets / financial_net_worth: FINANCIAL-ONLY view
(excludes hard assets like real estate, vehicles, jewelry).
These differ from the consumer-facing net worth and balance
sheet figures, which *include* hard assets and match the
dashboard.
- Plus: total_liabilities, annual income/expenses
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly, idempotent, non-destructive), so the bar is lower. The description nevertheless adds real behavioral context: edge cases where metrics are None (no debt accounts with minimum_payment, no mortgage accounts) and the warning that the balance-sheet DTI legitimately exceeds 1.0. It stops short of describing formatting or aggregation caveats.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded and the long metric list is well organized with clear labels. It is verbose for a parameterless read tool and 'Returns pre-computed ratios.' slightly restates the opening sentence, but nearly every line carries differentiating information about the metrics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden of explaining return values, and it does so exhaustively – naming each metric, its formula, its unit, and its interpretation. For a no-parameter, read-only analytics tool this leaves no significant gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4 per the rubric. There is nothing for the description to compensate for on the input side.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('get key financial ratios computed from the user's data') and clarifies the scope of what is returned. It implicitly distinguishes itself from siblings like get_net_worth and get_balance_sheet by noting the financial-only view excludes hard assets, but it never names an alternative tool outright, so routing still requires inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied through metric descriptions ('suited to net-worth analysis', 'measures housing affordability'), which hint at when each ratio matters. But there is no explicit statement of when to call this tool versus get_net_worth, get_balance_sheet, or get_income_statement, and no exclusions or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_forum_postForum: Get forum postARead-onlyIdempotentInspect
Get a single forum post, optionally with its replies.
Args:
post_id: The post ID.
include_replies: When True (default), include the post's replies.
Returns:
``{"post": {...}, "replies": [{id, post_id, author_id, content,
created_at, updated_at}]}`` (``replies`` omitted when
include_replies is False), or ``{"error": "Forum post not found"}``
(also returned for a staff-only post viewed by a non-staff user).
| Name | Required | Description | Default |
|---|---|---|---|
| post_id | Yes | ||
| include_replies | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive. The description adds meaningful behavioral detail: the exact return structure, that replies are omitted when include_replies=False, and that a not-found error is also returned for staff-only posts viewed by non-staff users. This access-control nuance is valuable beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-organized: a leading one-sentence summary, followed by Args and Returns sections. No filler, front-loaded, and every sentence adds useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read tool with 2 params and no output schema, the description covers purpose, parameters, return shape, and error cases, including permission-related errors. The only minor gap is the opaque {'post': {...}} object, which leaves the post's fields unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description compensates by explaining both parameters: post_id as the post ID and include_replies with its default and effect. This goes beyond bare property names and gives the agent enough to invoke correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the exact verb and resource: 'Get a single forum post, optionally with its replies.' This clearly distinguishes from sibling tools like list_forum_posts (plural) and create/update/delete_forum_post by specifying 'single' and the optional replies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implied usage is clear (fetch a single post, optionally with replies), but the description never names alternatives or conditions for choosing this over list_forum_posts or other forum tools. The guidance is embedded in the purpose, not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_gamification_summaryGamification: Get gamification summaryARead-onlyIdempotentInspect
Get the user's gamification summary including badges, streaks, level, and health score.
Returns:
Combined gamification data: badges, streaks, level info, and financial health score
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, covering the safety profile. The description adds the return content (badges, streaks, level, health score) but does not disclose any additional behavioral traits such as response format, performance implications, or authentication requirements. With annotations covering the main safety aspects, this is adequate but not rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no fluff. It front-loads the purpose and then clarifies the return data. Every word contributes to understanding the tool's function. It is appropriately sized for a zero-parameter read operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with no parameters and no output schema, the description is mostly complete. It specifies what data is returned, though it could be more explicit about the exact structure of the response (e.g., field names) or edge cases like empty gamification data. However, given the simplicity, it meets the minimum bar for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema provides no parameter documentation. Per the calibration, the baseline for 0 params is 4, and the description doesn't need to explain parameters. It correctly focuses on the output content, which is the only relevant semantic information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'get' and the resource 'gamification summary', listing the specific components (badges, streaks, level, health score). This distinguishes it from sibling tools that focus on individual aspects, such as get_budget_streaks or get_financial_health_score, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. There are sibling tools like get_points_and_level, get_budget_streaks, and get_financial_health_score that overlap in content. An agent would have to infer that this is the aggregated view without any explicit recommendation or differentiation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_goal_planGoals: Get goal planARead-onlyIdempotentInspect
Fetch one Visions & Goals plan in full: SMART clarifications
(Specific / Measurable / Achievable / Relevant / Time-bound answers)
and action plans with their steps.
Args:
vng_goal_id: The plan id (also stored in a linked zoninga
goal's ``vng_goal_id`` field).
Returns:
{"goal": {...clarifications, action_plans...}} or an error dict.
| Name | Required | Description | Default |
|---|---|---|---|
| vng_goal_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, openWorld, so safety is covered. The description adds genuinely useful behavioral context: the full shape of the returned payload (goal, clarifications, action_plans) and the failure mode ('or an error dict').
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Short, front-loaded with the core purpose, then Args and Returns. Every line earns its place; no filler or repetition of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and 0% parameter coverage, the description supplies both the return structure and the argument's provenance, which is what an agent needs. Only minor gaps remain (no pagination/error detail, a typo in 'zoninga').
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load, and it does: it explains the argument is the plan id and where to find it (a linked zoninga goal's vng_goal_id field). That is meaning well beyond the bare 'integer' in the schema, though it doesn't state format constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Fetch) and resource (one Visions & Goals plan) plus its contents: SMART clarifications and action plans with steps. It implicitly contrasts 'one ... plan in full' with the sibling list_goal_plans, though it never names that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to use this versus list_goal_plans, get_goal_projection, or update_goal. The 'one plan in full' phrasing hints at a read-one use case but no explicit context or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_goal_projectionAnalytics: Get goal projectionARead-onlyIdempotentInspect
Project when a financial goal will be reached at the current pace.
Sums contributions over the last 90 days and divides by the number of
months those contributions actually span (clamped to 1-3 months — a
goal funded for only a few weeks is NOT divided across a full 3 months),
then estimates months remaining and a projected completion date.
``pace_window_months`` reports that divisor; a value < 3 means the
projection rests on a short, preliminary window. If the goal has a
target date, also calculates the required monthly contribution to meet
that deadline.
Args:
goal_id: The goal's ID
Returns:
Current progress, avg monthly contribution, pace_window_months,
projected completion date, months remaining, and required monthly
to hit deadline.
| Name | Required | Description | Default |
|---|---|---|---|
| goal_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds significant behavioral detail beyond the annotations (readOnlyHint, idempotentHint, destructiveHint). It explains the 90-day summing, the clamping of the divisor to 1-3 months, and the meaning of pace_window_months. This goes beyond the safety profile already declared by annotations. No contradiction exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a clear opening statement, detailed calculation notes, and separate Args/Returns sections. It front-loads the core purpose and includes useful details like the clamping rule. It is slightly verbose but each sentence contributes value, so it earns a 4.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only analytic tool with a single integer parameter, the description is complete. It explains the calculation, the meaning of pace_window_months, and lists all return values (progress, average contribution, etc.), even though there is no output schema. An agent has everything needed to call it correctly and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, goal_id, is described as 'The goal's ID' – which restates the name and offers no additional meaning beyond the schema. With 0% schema description coverage, the description should compensate, but it does not explain how to obtain the ID or its format. This is minimal but at least acknowledges the parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Project when a financial goal will be reached at the current pace.' It uses a specific verb (project) and resource (financial goal), and the title 'Analytics: Get goal projection' reinforces this. It distinguishes itself from sibling tools like get_goal_plan (which likely provides a plan) by focusing on projection and pace calculation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool (when you need a projection based on current pace) and explains the calculation logic. However, it does not explicitly mention alternatives or when not to use it, though the purpose is unambiguous. Since it gives clear context without exclusions, it earns a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_income_statementReports: Get income statementARead-onlyIdempotentInspect
Get an income statement (profit & loss) for a time period.
The window spans exactly ``months`` calendar months back from today.
When that predates the user's first transaction, the start is clamped
to the first-activity date and ``period.clamped_to_first_activity`` is
true (``period.requested_start_date`` echoes the un-clamped start) —
so a "6 months" label backed by only 3 months of real data is
detectable, matching the other analytics tools (rule 35).
Args:
months: Number of months to cover (default: 6). Values <= 0 return
a validation error. Values > 24 are clamped to a 24-month window
but ``period.requested_months`` / ``period.clamped_months`` echo
the original request so the clamp is detectable.
Returns:
Income and expense breakdown by category with totals, plus a
``period`` block with start/end/months + the clamp metadata
(requested_months, clamped_months, requested_start_date,
clamped_to_first_activity).
| Name | Required | Description | Default |
|---|---|---|---|
| months | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare read-only/idempotent/non-destructive, and the description adds substantial behavior beyond them: calendar-month window semantics, clamping to first-activity date with detectable metadata, and the <=0 / >24 validation/clamp rules. This is rich behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose, then behavior, then Args/Returns. Well-organized but somewhat verbose with repeated clamp-detection narration across the period and months paragraphs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the Returns block is needed and present, describing both the income/expense breakdown and the period metadata fields. Covers everything an agent needs to call and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and only a bare integer 'months' with default exists, so the description must compensate fully. It does: describes default, boundary behavior (<=0 error, >24 clamped), and the echo fields that make clamping detectable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Get an income statement (profit & loss)') and names the exact time window ('months' calendar months back). Distinguishable from siblings like get_balance_sheet and get_cash_flow by the explicit P&L framing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage (retrieving period analytics) and references 'rule 35' and 'other analytics tools', but never states when to choose this over get_period_summary, get_spending_summary, or get_cash_flow. No explicit when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_leaderboardFamily: Get family leaderboardARead-onlyIdempotentInspect
Get the family gamification leaderboard.
Shows rankings by total points, badges earned, and streaks
for all members of your family group.
Returns:
Ranked list of family members with gamification stats.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive behavior. The description adds scope and output context by specifying that it ranks all family members by points, badges, and streaks and returns a ranked list, which is valuable since no output schema is provided.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short, front-loaded sentences with no filler: what it gets, what it shows, and what it returns. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only getter, the description covers the object, scope, content, and return shape despite the lack of an output schema. It is complete enough, though it could be stronger by explicitly distinguishing itself from nearby gamification or leaderboard siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, so the schema fully covers the argument surface. The description reinforces that the call takes no options and always produces a family-wide ranking, which exceeds the baseline for a zero-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Get the family gamification leaderboard,' then details the content ('rankings by total points, badges earned, and streaks') and scope ('all members of your family group'). This is enough to distinguish it from individual gamification and summary tools in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes clear that this tool is for a family-wide leaderboard, so an agent can infer when to call it. However, it gives no explicit when-to-use guidance and does not contrast itself with overlapping sibling tools such as get_gamification_summary, get_points_and_level, or list_badges.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_merchant_spendingAnalytics: Get merchant spendingARead-onlyIdempotentInspect
Get total spending at a specific merchant or matching a description.
Answers questions like "how much have I spent at Amazon/Starbucks?"
Searches transaction descriptions (case-insensitive partial match).
Args:
search_term: Merchant name or description to search for
start_date: YYYY-MM-DD (defaults to 12 months ago)
end_date: YYYY-MM-DD (defaults to today)
Returns:
Total spent, transaction count, average per transaction,
first/last dates, and monthly breakdown.
| Name | Required | Description | Default |
|---|---|---|---|
| end_date | No | ||
| start_date | No | ||
| search_term | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so the safety profile is covered. The description adds genuinely behavioral detail beyond that: the search is a case-insensitive partial match on transaction descriptions, and the date window defaults to the trailing 12 months when omitted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then structured Args/Returns sections. Slightly padded by the example-question sentence, but every block (scope, matching mechanism, parameters, return shape) earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema present, the description compensates by enumerating the return shape (total spent, transaction count, average, first/last dates, monthly breakdown). Combined with full parameter documentation and the matching mechanism, nothing needed to invoke or interpret the call is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden and does so: it explains search_term as a merchant name/description matched case-insensitively, gives the YYYY-MM-DD format for both date params, and supplies the default behavior for each.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource with explicit scope: 'Get total spending at a specific merchant or matching a description.' The merchant-scoped framing distinguishes it from broader siblings like get_spending_summary and get_category_trend, so an agent can route without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The example questions ('how much have I spent at Amazon/Starbucks?') give clear context for when this tool applies. However, it names no alternatives or exclusions — nothing tells the agent why to pick this over get_spending_summary or list_transactions for the same question.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_monitoring_planAgent: Get monitoring planARead-onlyIdempotentInspect
Get the user's agent monitoring plan — what the sentinel watches for
them (bill coverage, low balance, idle surplus, cashflow shortfall,
anomalies) plus the check-in cadence.
While the plan is DRAFT (not yet approved/activated), the sentinel does
nothing — activation is the user's opt-in to hands-off monitoring.
Args:
refresh_snapshot: Also run a fresh survey of the user's finances
(upcoming bills, income cadence, available funds) and include it
as "snapshot" — useful when helping the user set up or review
their plan.
Returns:
{plan: {...}, snapshot?: {...}}
| Name | Required | Description | Default |
|---|---|---|---|
| refresh_snapshot | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is known. The description adds valuable behavioral context beyond annotations: while the plan is DRAFT, the sentinel does nothing, and activation is the user's opt-in. It also clarifies the behavior of refresh_snapshot as running a fresh survey, which is a behavioral trait not captured by annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a clear lead sentence, a behavioral note about DRAFT plans, then Args and Returns sections. It is appropriately sized—every section earns its place and the most important information is front-loaded. No filler words or repetition of annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has one optional parameter and no output schema, the description is nearly complete. The Returns section gives the output shape {plan: {...}, snapshot?: {...}} and the plan content is described in the first sentence. The internal structure of 'plan' is not detailed, but because the tool is read-only and the plan's content is explained, nothing critical is missing for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for refresh_snapshot. The Args section fully explains what the parameter does (runs a fresh survey of finances), what it includes (upcoming bills, income cadence, available funds), how it affects output (adds a 'snapshot' field), and when it's useful. This exceeds what a schema-only description would provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get') and resource ('the user's agent monitoring plan'), and elaborates on exactly what the plan contains: bill coverage, low balance, idle surplus, cashflow shortfall, anomalies, and check-in cadence. This clearly distinguishes it from any sibling like update_monitoring_plan or get_setup_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context is clear: this tool retrieves the monitoring plan and optionally refreshes a snapshot. The refresh_snapshot parameter includes the use case 'useful when helping the user set up or review their plan.' It does not explicitly name alternative tools or state when not to use it, so it falls short of a 5 but is clearly context-rich.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_net_worthAccounts: Get net worthARead-onlyIdempotentInspect
Calculate the user's current net worth.
Includes financial accounts AND hard assets (real estate, vehicles, etc.)
to match the web UI dashboard calculation.
Returns:
Net worth breakdown with assets, liabilities, hard assets, and totals
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering safety. The description adds meaningful context: it explicitly includes hard assets and states it matches the dashboard calculation, clarifying the exact calculation scope. No contradictions with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each serving a purpose: the primary action, the scope inclusion, and the return structure. Front-loaded with the verb and resource, no redundant phrasing, and easily scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and no output schema, the description is fully adequate. It defines what the tool does, what is included, and what the response will contain. Combined with the rich annotations, an agent can invoke it correctly without further clarification.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema provides no burden. The description explains what the return value contains (assets, liabilities, hard assets, totals), which is helpful given there is no output schema. This satisfies the baseline for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool calculates the user's current net worth, specifies it includes financial accounts and hard assets, and explicitly ties it to the web UI dashboard calculation. This differentiates it from siblings like get_net_worth_history (historical) and get_balance_sheet (likely a broader or different financial statement).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by scoping the calculation to match the web UI dashboard, but it does not explicitly state when to use this tool versus alternatives such as get_net_worth_history or get_balance_sheet. There is no direct 'use this when' or 'instead of' guidance, leaving the agent to infer from the scope described.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_net_worth_historyAnalytics: Get net worth historyARead-onlyIdempotentInspect
Get net worth at the end of each month over time.
Reconstructs historical net worth by reversing transactions
from current balances. Shows assets, liabilities, net worth,
and month-over-month change for each period.
Args:
months: Number of months of history (default 12). Values <= 0 and
values > 1200 return a validation error; values > 24 (but <= 1200)
are capped to 24 but ``requested_months`` echoes the original
request, which makes the clamp detectable.
Returns:
Monthly net worth snapshots with change amounts and percentages,
plus ``requested_months`` and ``clamped_to_first_activity``. The
current (in-progress) month row is dated today with is_partial=true.
| Name | Required | Description | Default |
|---|---|---|---|
| months | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly, idempotent, non-destructive, and closed-world behavior. The description adds substantial context beyond that: it discloses the reconstruction method (reversing transactions from current balances), the validation and clamping behavior for months (including the detectable requested_months echo), and the partial-month row semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded and well-structured with explicit Args and Returns sections. It repeats the schema's default of 12 and could be slightly tighter, but nearly every sentence adds useful information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description explains the return shape (assets, liabilities, net worth, change amounts and percentages, requested_months, clamped_to_first_activity, is_partial) and the derivation caveat, giving an agent enough to call and interpret the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the full burden, and it does: default value, lower-bound and upper-bound validation errors, the 24-month cap, and the requested_months echo that makes clamping detectable. This is comprehensive parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Get net worth at the end of each month over time') and scopes it to historical monthly snapshots by explaining reconstruction from transactions. It does not explicitly name or contrast with the sibling get_net_worth, so it misses full sibling differentiation, but the temporal scope makes the distinction clear enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool's use for historical monthly net worth, but it never states when to choose it over get_net_worth or other analytics tools, and provides no when-not guidance. Usage is inferable from the purpose, which fits the 'implied usage' level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_period_summaryAnalytics: Get period summaryARead-onlyIdempotentInspect
Compare income, expenses, net, and savings rate across periods.
Covers month-over-month, quarter-over-quarter, and year-over-year
comparisons. Each period includes change_from_previous_pct.
Phase 114 — periods that end before the user's first transaction
are dropped, so a brand-new user asking for ``num_periods=6`` may
see fewer rows. Without this filter the leading zero-padding
rows produced misleading change-from-previous percentages (e.g.
-29.7% comparing the user's first real month to a prior $0
month). Response carries ``clamped_to_first_activity`` (bool)
and ``requested_periods``. When clamped, the returned rows are the
full extent of the user's history (e.g. 2 months of data); no
rows exist for the dropped periods.
Args:
period_type: 'month', 'quarter', or 'year'
num_periods: How many periods to include (default 6)
compare_prior_year: Include same period from prior year
Returns:
List of periods with income, expenses, net, savings_rate_pct,
and change percentages, plus the two clamping metadata fields.
| Name | Required | Description | Default |
|---|---|---|---|
| num_periods | No | ||
| period_type | No | month | |
| compare_prior_year | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the read-only/idempotent safety profile, but the description goes well beyond them by disclosing the clamping behavior: periods ending before first activity are dropped, returned rows may be fewer than num_periods, and two metadata fields (clamped_to_first_activity, requested_periods) signal this. That edge-case disclosure is exactly the extra value an agent needs to interpret results correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then structured Args/Returns sections that are useful given there is no output schema. The 'Phase 114' changelog framing is internal jargon, but the underlying clamping explanation earns its space because it prevents misinterpretation of change percentages.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing the returned fields (income, expenses, net, savings_rate_pct, change_from_previous_pct) and the two clamping metadata fields. Combined with the annotations, an agent has everything needed to call and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning, and it does: it enumerates the allowed period_type values ('month', 'quarter', 'year') that the schema omits, plus the num_periods default and the compare_prior_year flag. Only minor detail (e.g., valid range for num_periods) is missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — compares income, expenses, net, and savings rate across month/quarter/year periods — so the agent knows exactly what it does. It does not, however, explicitly distinguish itself from the many adjacent analytics tools (get_savings_rate_trend, get_spending_summary, get_income_statement).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The mention of month-over-month, quarter-over-quarter, and year-over-year comparisons implies the use case (trend/period comparison), giving clear implied usage. There is no explicit when-to-use guidance, no named alternatives, and no stated exclusions versus the sibling trend/summary tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_plaid_itemPlaid: Get linked-bank detailsARead-onlyIdempotentInspect
Get details for a single Plaid item, including the accounts it owns.
Args:
item_id: The PlaidItem.id (not the upstream Plaid item_id)
Returns:
Item metadata + a nested list of linked accounts
(id, name, type, balance, plaid_last_balance_update).
| Name | Required | Description | Default |
|---|---|---|---|
| item_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, covering safety. The description adds context about the return payload (metadata + nested accounts) and the id caveat, but does not disclose any side effects or additional behavioral traits. This is adequate but not rich, hence a 3.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured with Args and Returns sections. It leads with the purpose, then clarifies the parameter, then lists the return structure. Every sentence earns its place with no repetition or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with a single parameter, the description is complete: it specifies the purpose, the exact id to use, and the shape of the response. Annotations cover safety, and sibling tools are distinguishable. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no description for item_id (0% coverage), but the description fully compensates by explaining that item_id is the PlaidItem.id, not the upstream ID. This is essential information for correct invocation and goes beyond what the schema alone could convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'Get details for a single Plaid item, including the accounts it owns.' It distinguishes itself from list_plaid_items by specifying 'single' item, and clarifies that item_id is the internal PlaidItem.id, not upstream. This leaves no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies its usage context: use for a single item's detailed view; use list_plaid_items when enumerating items. It also notes the id distinction, which prevents misuse. However, it does not explicitly name sibling alternatives or state when not to use it, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_points_and_levelGamification: Get points and levelARead-onlyIdempotentInspect
Get the user's XP points, level, and recent point history.
Returns:
Level info, total points, progress to next level, and recent 20 point entries
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint), so the description adds value by detailing the return contents: level info, total points, progress to next level, and a 20-entry point history cap. This is useful behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no filler. The first sentence states the tool's purpose, and the second concisely enumerates the returned data. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, zero-parameter read-only getter with no output schema, the description fully explains what the agent can expect: level info, total points, progress, and the specific 20-entry history limit. No critical call-time information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and schema coverage is 100%, so there are no parameter semantics for the description to clarify. Per the baseline for zero-parameter tools, this is a strong score; the description need not discuss parameters it does not have.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: getting the user's XP points, level, and recent point history. It clearly identifies the resource, but it does not distinguish itself from the closely related sibling get_gamification_summary, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when the tool is useful—whenever an agent needs XP, level, or point history—but it does not explicitly state when to prefer it over alternatives like get_gamification_summary or get_leaderboard. Usage context is clear from the resource name but exclusions are absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_savings_rate_trendAnalytics: Get savings-rate trendARead-onlyIdempotentInspect
Monthly savings rate over time.
Shows income, expenses, net, and savings rate percentage for
each month. Answers questions like "is my savings rate improving?"
Phase 114 — months that predate the user's first transaction
are dropped (same logic as the period summary). A brand-new
user asking for a 12-month trend gets only the months they
actually have data for, not 10 leading $0 rows. Response
carries ``clamped_to_first_activity`` and ``requested_months``.
Args:
months: Number of months (default 12)
Returns:
Monthly savings rate data, plus the two clamping metadata
fields.
| Name | Required | Description | Default |
|---|---|---|---|
| months | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), and the description goes further by disclosing the Phase 114 clamping behavior, the fact that leading $0 months are dropped, and the two metadata fields (clamped_to_first_activity, requested_months) returned. This is genuinely non-obvious behavior an agent could not get from the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose sentence is front-loaded and the return/metadata notes earn their place, but the 'Phase 114 —' internal roadmap reference is developer-facing noise that adds little for an agent. Conventional Args/Returns block is slightly redundant with the prose above it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read tool with no output schema, the description covers what it returns, the metric breakdown, and the clamping metadata, which is enough to call and interpret it. Missing only pagination/size caveats and tighter parameter bounds.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description supplies only 'months: Number of months (default 12)'. It adds conceptual meaning via the 12-month example and clamping narrative, but does not state bounds, allowed ranges, or edge behavior for extreme values, so compensation for the coverage gap is partial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific metric (savings rate) and granularity (monthly over time), and clarifies sub-fields (income, expenses, net, savings rate percentage). It also ties itself to the intent 'is my savings rate improving?'. It references the period summary for the clamping logic but does not strongly disambiguate against close analytics siblings like get_financial_ratios or get_cash_flow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied via the example question ('is my savings rate improving?'), which suggests the analytics intent. There is no explicit when-to-use/when-not guidance and no named alternative for related trend questions, so selection against the many sibling analytics tools is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_scenarioScenarios: Get a what-if scenario with its projected impactARead-onlyIdempotentInspect
Get one scenario, its hypothetical changes, and the projected impact.
Runs the debt-paydown engine on the user's real debts adjusted by this
scenario's changes and returns the deltas vs their current plan:
payoff-date change, total-interest change, monthly-payment change,
net-worth change, and the best paydown strategy under the scenario.
Args:
scenario_id: The scenario to analyze.
Returns:
{scenario: {...}, items: [...], analysis: {best_strategy,
interest_delta, monthly_payment_delta, net_worth_delta,
baseline_payoff_date, scenario_payoff_date, ...}} or
{"error": "..."} if not found.
| Name | Required | Description | Default |
|---|---|---|---|
| scenario_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark readOnlyHint, idempotentHint, and destructiveHint, lowering the burden. The description adds valuable context about running the debt-paydown engine, computing deltas against the current plan, and returning an error when the scenario is not found. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a clear one-sentence summary followed by a brief explanation and Args/Returns sections. It is well-structured and avoids waste, though the Returns section's '...' placeholder is slightly untidy and could be tightened.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only tool without an output schema, the description covers the main return shape, including several analysis fields and the error case. However, 'items: [...]', 'scenario: {...}', and the trailing '...' leave parts of the return undocumented, making it adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It offers only 'scenario_id: The scenario to analyze,' which essentially restates the parameter name and adds no guidance on sourcing the ID, allowed values, or constraints. This is insufficient for a parameter left undocumented by the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Get one scenario, its hypothetical changes, and the projected impact.' It further clarifies the tool's unique function by describing the debt-paydown engine and the return deltas, which clearly distinguishes it from list_scenarios, create_scenario, update_scenario, and delete_scenario.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use when a user needs a single scenario's projected impact and best strategy, but it never names alternatives or gives exclusion criteria. An agent must infer that list_scenarios is for enumeration and commit_scenario for applying, rather than being told explicitly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_seasons_summaryGamification: Get seasons summaryARead-onlyIdempotentInspect
Get seasonal challenge sets with user's progress.
Returns:
Active, upcoming, and past seasons with per-challenge progress
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the operation is read-only, idempotent, and non-destructive, so the description carries a light burden. It adds value by specifying the return composition (active/upcoming/past seasons and per-challenge progress), which is especially helpful because no output schema is provided.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short lines, front-loads the core action, and uses a compact Returns block rather than repeating the schema. No filler or redundant restatement of the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, annotation-covered getter, the description covers the request object and high-level response shape. It would be slightly more complete with an explicit relationship to get_gamification_summary/list_challenges, but nothing essential for invoking the tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero parameters, the schema already fully covers the input contract, and the description adds the only relevant semantic info: the result is the user's own progress. The baseline of 4 applies because there is nothing for the description to explain about parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Get seasonal challenge sets') and scopes the result to 'active, upcoming, and past seasons' with per-challenge progress. This clearly separates it from generic challenge-list or gamification-summary siblings without requiring schema inspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended context is implicit: use when a summary of seasonal challenges and the user's progress is needed. However, there is no explicit when-to-use statement or mention of alternatives such as get_gamification_summary or list_challenges, so an agent must infer routing from the name and return description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_setup_statusOnboarding: Get setup statusARead-onlyIdempotentInspect
Get the user's onboarding/setup progress — the agent's setup curriculum.
Reports what the user still needs to be "hands-off ready", computed
from live data. Most useful near the start of a new user's
conversation.
Finish line is "fast to hands-off": at least one account, an income
signal, and an ACTIVE daily-monitoring plan. Budgets and goals are
offered but optional.
Returns:
``is_new_user`` (no accounts yet), ``hands_off_ready`` (all required
steps done), ``next_step`` / ``next_step_label``,
``required_remaining`` (keys), the ordered ``steps`` (each
``{key, label, done, required, detail}``), ``completed_count`` /
``total_count``, and a ``guidance`` string describing the
recommended next setup step. When ``hands_off_ready`` is True,
required setup is complete.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safe read-only/idempotent profile, but the description adds real value beyond them: results are 'computed from live data', the completion criteria are defined (one account, an income signal, an ACTIVE daily-monitoring plan), and budgets/goals are flagged as optional. It could say more about latency of the live computation, but the added context is substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, and the rest is organized into scannable blocks. It is somewhat long, but the detail is content-bearing, especially the Returns block, which compensates for the missing output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter, read-only status tool with no output schema, the description covers everything an agent needs: what is reported, how completion is determined, what is optional, and the exact shape of the returned fields including the guidance string and next_step.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is no argument surface that needs additional semantic explanation in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource combination ('Get the user's onboarding/setup progress') and frames it as the agent's setup curriculum, computed from live data. This clearly distinguishes it from lookalike siblings such as get_user_profile or get_monitoring_plan, which report persistent state rather than onboarding completion.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear usage context: 'Most useful near the start of a new user's conversation.' However, it names no alternative tool and gives no explicit when-not guidance (e.g., once hands_off_ready is True, use X instead), so it stops short of the 5 tier.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_spending_anomaliesAnalytics: Get spending anomaliesARead-onlyIdempotentInspect
Detect categories with unusually high spending this month.
Compares current month spending to recent monthly averages.
Flags categories where spending exceeds the threshold multiplier.
Args:
months_lookback: Number of months for computing averages (default 3)
threshold: Multiplier threshold (default 2.0 = 2x the average)
Returns:
List of anomalies with category, current amount, average,
ratio, and severity (high/medium).
| Name | Required | Description | Default |
|---|---|---|---|
| threshold | No | ||
| months_lookback | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so safety is covered. The description adds value by explaining the comparison logic, the threshold interpretation, and the return fields (category, current amount, average, ratio, severity). This is meaningful behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured with clear sections for Args and Returns. Every sentence conveys necessary information—purpose, algorithm, parameter semantics, and output shape—with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only analytics tool with two optional parameters and no output schema, the description is quite complete. It documents both parameters with meanings and defaults, and describes the return list fields. Minor gaps remain (e.g., exact calendar-month definition, empty-result behavior), but these are not critical for calling the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the full burden. It explains months_lookback as 'Number of months for computing averages (default 3)' and threshold as 'Multiplier threshold (default 2.0 = 2x the average)', adding exactly the semantic meaning the schema lacks. The defaults are also stated in clear prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Detect categories with unusually high spending this month.' This clearly distinguishes it from sibling analytics tools like get_spending_summary or get_category_trend by focusing on anomaly detection rather than summaries or trends.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the use case explicit: detecting categories with unusually high spending compared to recent averages. It does not name alternatives or give exclusions, but the algorithm description ('compares current month spending to recent monthly averages') gives clear contextual guidance on when this tool is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_spending_summaryReports: Get spending summaryARead-onlyIdempotentInspect
Get spending summary grouped by category.
Args:
start_date: Start date in YYYY-MM-DD format (defaults to 30 days ago)
end_date: End date in YYYY-MM-DD format (defaults to today)
Returns:
Spending breakdown by category
| Name | Required | Description | Default |
|---|---|---|---|
| end_date | No | ||
| start_date | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds valuable behavioral context beyond annotations: the default date range (30 days ago to today) and the grouping by category, which are not visible from the schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured: a one-sentence purpose, followed by Args and Returns sections. Every sentence earns its place, and the purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only report with two optional parameters, the description covers the purpose, inputs, defaults, and return concept. The return description is terse ('Spending breakdown by category'), but with no output schema, slightly more field-level detail would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden for both parameters. It fully compensates by explaining the YYYY-MM-DD format and the defaults for start_date and end_date, adding meaning the bare schema does not provide.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb and resource: 'Get spending summary grouped by category.' It is unambiguous about what the tool does, though it does not explicitly distinguish itself from related sibling tools like get_cash_flow or get_period_summary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to prefer this tool over alternatives, and it does not state any exclusions or prerequisites. It documents parameter defaults, but this is parameter semantics rather than usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_standing_factsAgent: Get standing factsARead-onlyIdempotentInspect
The user's standing facts in one call (Phase 251): their SAVED debt
plan resolved to THIS month — `saved_plan.this_month.target` is the
debt receiving the extra payment, with `payment`, `minimum_payment`,
`extra`, and `next_payment_due_date` — plus `upcoming_bills` for the
next `days` (declared schedules, Plaid due dates, AND recurring
bills observed in the bank history — `is_inferred: true`),
scheduled income, `checking_available`, and the monitoring plan
status.
Relevant to questions about paying debts, what to pay this month,
bills, upcoming money, or spare cash. For example, a target `payment`
of $252.48 is a $202.48 `minimum_payment` plus a $50 `extra`. A
non-empty `upcoming_bills` means bills are due in the window.
Args:
days: look-ahead window for bills/income (default 14, max 90)
Returns:
{as_of, saved_plan | null, upcoming_bills, upcoming_bills_total,
scheduled_income_total, checking_available, monitoring}
| Name | Required | Description | Default |
|---|---|---|---|
| days | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds genuinely new context: bills may be inferred from bank history and are flagged with is_inferred=true, saved_plan can be null, and a non-empty upcoming_bills signals bills due in the window. It stops short of describing pagination or failure behavior (e.g., no bank connection).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core statement, then the field semantics, then Args and Returns blocks — a clean structure. The dollar-value example sentence is illustrative and earns its place, though the prose is somewhat long for a single-parameter read tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so: it lists the returned keys ({as_of, saved_plan | null, upcoming_bills, upcoming_bills_total, scheduled_income_total, checking_available, monitoring}) and explains the meaning of the key nested fields. Combined with the read-only annotations, an agent has everything needed to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% (the single 'days' property carries no description), so the description must compensate — and it does, documenting the look-ahead semantics plus the default (14) and max (90) that the schema omits entirely. Only one parameter and no format subtleties keep this from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete aggregation: 'The user's standing facts in one call' and enumerates the payload members (saved debt plan resolved to this month, upcoming bills, scheduled income, checking_available, monitoring plan status). An agent knows exactly what this returns. It doesn't explicitly name which sibling it supersedes (get_monitoring_plan, get_cash_flow, get_debt_paydown all overlap thematically), so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit triggering questions: 'paying debts, what to pay this month, bills, upcoming money, or spare cash,' plus a concrete example ($252.48 = $202.48 minimum + $50 extra). Clear context for when to reach for it, but no stated exclusions or named alternative tools for when this is the wrong call.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_subscription_statusSubscription: Get subscription statusARead-onlyIdempotentInspect
Get the user's subscription plan, limits, and trial status.
Returns:
``plan`` (stable plan_type: "free" / "premium_monthly" /
"premium_yearly"), ``is_premium``, ``is_trialing``, optional
``trial_end`` / ``trial_days_remaining``, and ``limits`` — the Free
caps (``max_accounts`` / ``max_transactions_per_month`` /
``max_goals`` / ``max_budgets`` / ``max_plaid_connections``). A
limit of ``0`` means unlimited (Premium plans).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false. The description adds valuable behavioral detail beyond annotations: the exact return shape, optional trial fields, stable plan_type values, and the important semantic that a limit of 0 means unlimited for Premium plans. This fully explains what the tool returns and how to interpret it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a clear 'Returns:' section, uses recognizable type names, and explains edge-case semantics concisely. Every sentence adds value with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given this is a parameterless read-only operation with annotations covering safety and an output schema absent, the description is complete. It describes all returned fields, their types, optionality, and meaning, leaving no ambiguity for an agent invoking the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the input schema has no burden. The description adds no parameter documentation because none is needed. Baseline 4 is appropriate for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb and resource ('Get the user's subscription plan, limits, and trial status') and clearly distinguishes this read-only status tool from subscription lifecycle tools like cancel_subscription, reactivate_subscription, and list_subscription_plans. It tells an agent exactly what information is available.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context is clear: use this when you need the current user's subscription plan, limits, or trial status. It does not explicitly list alternatives or say when not to use it, but the purpose is specific enough that an agent can infer the right situation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_support_ticketSupport: Get support ticketARead-onlyIdempotentInspect
Get one of the user's own support tickets plus its full message thread.
API docs: https://zoninga.com/ai/#support-tickets
Args:
ticket_id: The ticket's ID.
Returns:
``{"ticket": {...}, "messages": [{id, message, is_staff_reply,
sender, created_at}]}`` ordered oldest-first, or
``{"error": "Support ticket not found"}`` for a ticket that
doesn't exist or isn't the caller's (no existence leak).
| Name | Required | Description | Default |
|---|---|---|---|
| ticket_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, so safety is covered. The description adds genuinely useful context beyond that: the thread is returned ordered oldest-first, and a missing/foreign ticket yields {"error": "Support ticket not found"} with an explicit no-existence-leak note, which is a meaningful security behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose sentence, then compact Args/Returns sections. The docstring formatting is slightly verbose but every element (scope, return shape, error case) earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so by giving the concrete response shape and the error variant. It is nearly complete for a one-parameter read tool, with the only gap being lack of routing to sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the only description of the parameter is 'ticket_id: The ticket's ID.', which adds little beyond the schema name. For a single, self-evident integer identifier this is adequate but does not compensate for the coverage gap (e.g., no hint on where the ID comes from, such as list_support_tickets).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Get) and resource (support ticket) with an important scope qualifier: 'one of the user's own', plus the addition of the full message thread. This clearly separates it from a bare list operation, though it never names the sibling tools (list_support_tickets, reply_to_support_ticket) to route the agent explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance or alternatives. The agent must infer that this is the single-ticket fetch (taking a ticket_id) versus list_support_tickets, and there is no mention of prerequisites or when NOT to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_uncategorized_summaryAnalytics: Get uncategorized summaryARead-onlyIdempotentInspect
Get count and total of uncategorized transactions.
Identifies transactions that need categorization.
Phase 115 — also returns a per-type breakdown (``by_type``) and
an ``actionable_count``/``actionable_total`` pair (deposits +
withdrawals only, excluding transfers). The income statement
excludes transfers entirely since they move no money in or out,
so an uncategorized TRANSFER counts in ``uncategorized_count``
here but is absent from the income statement's "Uncategorized"
line. ``actionable_count`` answers "what's blocking my income
statement from being clean?"; ``uncategorized_count`` is the
absolute completeness count.
Args:
start_date: YYYY-MM-DD (defaults to Jan 1 of current year)
end_date: YYYY-MM-DD (defaults to today)
Returns:
Count, total amount, percentage of all transactions that are
uncategorized, plus the by_type breakdown and the
actionable (income-statement-relevant) counts.
| Name | Required | Description | Default |
|---|---|---|---|
| end_date | No | ||
| start_date | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, idempotentHint, destructiveHint=false), so the description earns credit for going beyond them: it discloses the by_type breakdown, the two distinct counting semantics (actionable_count vs uncategorized_count), and the transfer-exclusion rule that explains why this tool's total can differ from the income statement. That is genuine behavioral context not derivable from annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose and organized into Args/Returns blocks that earn their place given 0% schema coverage and no output schema. The 'Phase 115 —' internal release-notes tag is the one piece of noise an agent gains nothing from.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only analytics tool with no output schema and no schema-level parameter docs, the description supplies formats, defaults, return contents, and the semantic distinction between the two counts. Only edge details (range inclusivity, empty-result behavior) are unaddressed, which is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden and does so: both parameters get a YYYY-MM-DD format and a concrete default (Jan 1 of current year / today). This is exactly the compensating detail the empty schema lacks, though it does not discuss inclusivity of the range boundaries.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Get count and total of uncategorized transactions,' plus a one-line intent gloss ('Identifies transactions that need categorization'). It also implicitly separates itself from the income-statement tool by explaining that uncategorized transfers count here but not there. It never names a sibling alternative directly, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied rather than stated – the 'what's blocking my income statement from being clean?' framing suggests the agent reaches for this when auditing data quality, but no explicit when-to-use/when-not clause or named alternative tool is given. An agent must infer the trigger from the analytics sibling set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_unread_notification_countNotifications: Get unread notification countARead-onlyIdempotentInspect
Get the count of unread notifications (for badge display).
Returns:
Unread notification count.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds modest context by framing the result as badge-oriented and restating that it returns a count, but it adds no deeper behavioral nuance such as counting scope or stale-data behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short lines with no filler. It front-loads the purpose and adds a returns note, both of which are directly relevant to an agent deciding whether to call this tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only count tool with rich annotations, the description is nearly complete. It states the purpose, the badge-display use case, and the return value; specifying that the count is an integer would be a minor enhancement.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the description carries little semantic burden. There is nothing about parameters that needs elaboration, and the schema already covers the empty parameter set.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Get the count of unread notifications' for badge display. This clearly distinguishes it from sibling tools like list_notifications, which return notification records, and mark_notification_read, which mutates state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(for badge display)' provides clear context for when this count-focused tool should be used. It does not explicitly name alternatives or exclusions, but the zero-parameter count-only design makes the usage context apparent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_user_preferencesProfile: Get user preferencesARead-onlyIdempotentInspect
Get the user's display, notification, privacy, and debt-plan preferences.
Returns:
Theme, currency, notification toggles, privacy settings, and the
saved debt-paydown plan (debt_strategy + debt_extra_monthly_payment).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=true, idempotentHint=true, and destructiveHint=false, fully covering the safety profile. The description adds useful context by enumerating the returned preference categories and fields, but reveals no additional behavioral traits beyond what the annotations and return list imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two compact lines: a clear one-sentence purpose followed by a structured return-value list. It is front-loaded, contains no filler, and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only tool with no output schema, the description provides a complete contract: it names the preferences retrieved and the exact return fields. Nothing essential is missing for an agent to decide whether and how to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description appropriately focuses on what the returned data contains rather than input semantics, which is exactly what an agent needs for a parameterless getter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves the user's preferences, listing specific domains (display, notification, privacy, debt-plan) and the exact return fields. This is a clear verb+resource with enough specificity to distinguish it from related getters like get_user_profile, though it does not explicitly name the sibling it differs from.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied through the stated purpose: call this when you need the user's display, notification, privacy, or debt-plan preferences. However, there is no explicit guidance about when not to use it or which alternative to choose, particularly among the many sibling get_* tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_user_profileProfile: Get user profileARead-onlyIdempotentInspect
Get the authenticated user's profile information.
Identifies which account the session is operating on. Returns the
user's email, name, and subscription plan.
Returns:
Dictionary with user identity and account details
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish the safe read profile (readOnlyHint, idempotentHint, non-destructive, closed-world). The description adds value beyond them by disclosing the returned content (email, name, subscription plan) in the absence of any output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded with the core action. The trailing 'Returns: Dictionary with user identity and account details' mildly restates the preceding sentence about return fields, costing a little efficiency but not much.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, side-effect-free read with no output schema, the description covers identity scope and return contents sufficiently for an agent to call it correctly. No notable gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline for a parameterless tool applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Get the authenticated user's profile information') and clarifies scope as the session's own account, which distinguishes it from sibling tools like update_user_profile or get_user_preferences. It does not explicitly name a sibling to route away from, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Identifies which account the session is operating on' implies the use case (confirming identity/context), but there is no explicit when-to-use, when-not-to-use, or pointer to alternatives such as get_user_preferences. Usage is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_weekly_recapsGamification: Get weekly recapsARead-onlyIdempotentInspect
Get recent weekly financial recaps (Monday summaries).
Args:
limit: Maximum number of recaps to return (default: 5, max: 20)
Returns:
List of weekly recap summaries
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds recency, Monday scope, and limit behavior, but it doesn't discuss ordering, authentication needs, or what fields are included in each summary. This is adequate but not deeply transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: a clear one-line purpose followed by clear Args and Returns sections. Every sentence serves a purpose; there is no filler or repetition of schema fields beyond necessary semantics.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter, read-only list tool, the description supplies the resource type, recency, weekly cadence, and limit constraints. The only gap is the lack of detail on the internal fields of a weekly recap, which would be valuable since there is no output schema, but it doesn't block correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter meaning. It fully does: 'limit: Maximum number of recaps to return (default: 5, max: 20)', which is precise and actionable for a single optional parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb and resource: 'Get recent weekly financial recaps (Monday summaries)'. It specifies the exact scope, is not a tautology, and is readily distinguishable from sibling gamification tools like get_gamification_summary or get_period_summary.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it — when recent weekly recap summaries are needed — and clarifies their cadence ('Monday summaries'). However, it does not explicitly mention alternatives or state when not to use it, leaving the agent to infer selection context from the name and wording.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_accountsAccounts: Import accounts from CSVAInspect
Bulk-create accounts from CSV text.
Parity with the web account-import flow, adapted for a chat agent: pass
the CSV as text (not a file). Columns are auto-detected (name + type are
required; balance, institution, interest_rate, credit_limit,
minimum_payment, term_months optional). Rows that can't be parsed are
skipped and reported; valid rows still import. Created atomically.
Account type accepts friendly strings (checking, savings, cash, credit
card, loan, mortgage, student loan, investment, other asset, other
debt). Respects the Free-plan account cap — if existing + new accounts
would exceed it, the whole import is rejected before any insert.
Args:
csv_text: CSV content as a string (header row + account rows). Max
1000 rows.
skip_duplicates: When True (default), skip rows whose name matches
an existing active account. When False, a duplicate name still
imports (creating a second account with that name).
Returns:
``{"success": True, "imported_count": N, "skipped_duplicates": K,
"error_count": M, "errors": [...up to 20...], "account_names":
[...]}`` on success, or ``{"error": "..."}`` (empty/invalid CSV, no
valid rows, all duplicates, or account-limit exceeded).
| Name | Required | Description | Default |
|---|---|---|---|
| csv_text | Yes | ||
| skip_duplicates | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description goes well beyond the sparse annotations. It discloses partial-failure behavior (invalid rows are skipped and reported while valid rows import), atomicity (created atomically), the Free-plan account cap enforcement (whole import rejected before any insert), and the exact return/error shape. This gives an agent a thorough picture of what will happen.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured with clear paragraphs and an Args/Returns block. It is longer than minimal but every sentence carries useful information; the core purpose is front-loaded. A minor deduction because the return-value section is somewhat verbose, though it is valuable with no output schema present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without an output schema, the description specifies the success and error return formats, input constraints (max 1000 rows, required/optional columns), duplicate handling, and the account cap behavior. For a bulk-import tool with both validation and partial-failure semantics, this is complete enough for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description fully compensates: it explains csv_text as CSV string with header row, max 1000 rows, and describes skip_duplicates behavior with both True and False outcomes. It also explains the auto-detected columns and required vs optional fields, adding meaning that the schema completely lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Bulk-create accounts from CSV text' – a specific verb, resource, and method that clearly distinguishes this from single-account creation via create_account and from import_transactions. The title and first sentence together make the tool's scope immediately obvious.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly establishes the context: this is for bulk-importing accounts from CSV text, adapted from the web flow for a chat agent. It explains how to pass CSV as text and what fields are auto-detected, but it does not explicitly name alternatives or say when not to use it (e.g., for single accounts or transaction imports).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
import_transactionsTransactions: Import transactions from CSVAInspect
Bulk-import transactions into one account from CSV text.
Parity with the REST ``/api/v1/data/import_transactions/`` upload,
adapted for a chat agent: pass the CSV as text (not a file). All rows
are created atomically under a row lock and the account balance is
updated in one pass.
CSV format: a header row then ``date,description,amount,type``.
- date: YYYY-MM-DD (also accepts MM/DD/YYYY or MM/DD/YY)
- amount: positive number ($ and commas are stripped)
- type: deposit/income/credit → deposit; withdrawal/expense/debit →
withdrawal; blank → ``default_type``. (No transfers — import each
side as a deposit/withdrawal.)
Rows that can't be parsed are skipped and reported in ``errors``;
valid rows still import.
On debt accounts the deposit/withdrawal meaning is inverted (a
``deposit`` is a charge that increases the debt,
a ``withdrawal`` is a payment that decreases it) — the balance math is
identical either way. Future-dated rows are created but not applied to
the balance until their date arrives (matching single-transaction
creation).
Args:
account_id: Target account (the caller's own, active).
csv_text: The CSV content as a string (header + rows). Max 1000
rows per call — split larger files.
default_type: Type for rows with a blank ``type`` column
("deposit" or "withdrawal", default "withdrawal").
Returns:
``{"success": True, "imported_count": N, "error_count": M,
"errors": [...up to 20...], "new_account_balance": ...,
"account_name": ...}`` on success, or ``{"error": "..."}`` (account
not found, empty/invalid CSV, monthly-limit exceeded, or no valid
rows).
| Name | Required | Description | Default |
|---|---|---|---|
| csv_text | Yes | ||
| account_id | Yes | ||
| default_type | No | withdrawal |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: all rows created atomically under a row lock, balance updated in one pass, partial-success semantics (bad rows skipped and reported while valid rows still import), debt-account sign inversion, and future-dated rows not applied until their date. Failure modes (account not found, empty/invalid CSV, monthly-limit exceeded, no valid rows) are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the one-line purpose before format details, and organized into CSV / behavior / Args / Returns blocks. It is long, and a couple of lines (e.g. the REST-parity note and repeated balance-math remark) could be trimmed, but nearly every sentence carries decision-relevant info.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description fully specifies the return envelope (success/imported_count/error_count/errors capped at 20/new_account_balance/account_name) and error shape, plus all edge-case behavior. An agent has everything needed to call and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the load — and it does: account_id is defined as the caller's own active account, csv_text is defined as header+rows with a 1000-row cap, and default_type is enumerated as deposit/withdrawal with its default and its role for blank type cells.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource+scope: 'Bulk-import transactions into one account from CSV text.' It is clearly distinguishable from siblings like create_transaction (single), import_accounts, and parse_statement, and it names the REST endpoint it mirrors.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong operational context: pass CSV as text (not a file) because it is chat-adapted, plus a hard 'Max 1000 rows per call — split larger files' constraint and a note to import transfer sides separately. It never explicitly routes the agent away from create_transaction or other import tools, so it stops short of full when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
join_challengeGamification: Join challengeCInspect
Join a challenge for the current period.
Args:
challenge_slug: The slug identifier of the challenge to join
Returns:
Confirmation with challenge details and period dates
| Name | Required | Description | Default |
|---|---|---|---|
| challenge_slug | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate this is a mutation (readOnlyHint=false) and not idempotent (idempotentHint=false). The description only says 'Join a challenge' and that it returns confirmation, but it does not disclose side effects (e.g., affecting points, requiring active membership, or that calling twice could cause errors). It adds minimal behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise, front-loaded with the core purpose, and efficiently lists the argument and return value. Every sentence serves a purpose with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description mentions the return type (confirmation with details and period dates), which is helpful given there is no output schema. However, it lacks usage guidance, side effects, and any relation to sibling gamification tools. For a single-parameter mutation tool, it is adequate but has clear gaps in when and how to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. The only parameter, challenge_slug, is described as 'The slug identifier of the challenge to join,' which essentially restates the parameter name without adding meaningful detail (e.g., that it must come from list_challenges or be an active challenge). This provides little additional value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Join a challenge for the current period.' This is a specific verb and resource, and it is distinct from sibling tools like list_challenges and leave_challenge. However, it does not explicitly differentiate itself from alternatives, so it misses the full mark for distinguishing from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention prerequisites like needing to call list_challenges first, nor does it indicate when not to use it (e.g., if the challenge is not open). No exclusions or contextual hints are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
leave_challengeGamification: Leave challengeAInspect
Leave an active challenge for the current period.
Removes the user from the challenge so it can be re-joined later (a
re-join re-counts the period's activity). Completed challenges are not
affected.
Args:
challenge_slug: The slug identifier of the challenge to leave
Returns:
Confirmation, or an error if the user wasn't in that challenge.
| Name | Required | Description | Default |
|---|---|---|---|
| challenge_slug | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the important behavioral nuance: leaving is reversible, re-joining re-counts the period's activity, and completed challenges are untouched. This goes well beyond the minimal annotations and tells the agent exactly what side effects to expect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: a clear one-line summary, behavioral details, args, and returns. Every sentence adds relevant information without unnecessary filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter mutation with no output schema, the description is complete: it explains what happens, when it applies, what the side effects are, and what the return will be. An agent has enough information to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero schema description coverage, the description must explain the parameter. It does identify challenge_slug as 'the slug identifier of the challenge to leave,' which provides basic meaning, but it offers no guidance on where to find the slug or how slug format should be validated.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Leave an active challenge for the current period.' It also clarifies scope by stating completed challenges are not affected, which distinguishes it from operations on completed or historical challenges.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly indicates this tool is for leaving an active challenge and notes that completed challenges are not affected. It does not explicitly name join_challenge as the alternative, but the context and sibling list make the intended usage clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
leave_family_groupFamily: Leave family groupADestructiveInspect
Voluntarily leave your current family group. The group owner can't
leave; the error points them to deleting the group on the Family page.
Returns:
Confirmation on success, or error dict
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, so the destructive nature is covered. The description adds useful behavioral details: the action is voluntary, the owner is blocked and will get an error directing them elsewhere. This goes beyond the structured annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences plus a return note; every sentence serves a purpose. The primary action is front-loaded, followed by the key exception and return type. No fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, single-action tool with destructive annotations, the description is complete: it defines the action, the owner limitation, and the return value. No output schema exists, but the return note covers what the agent needs.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has 0 parameters and schema coverage is 100%, so there is no parameter semantics to convey. The description does not need to compensate; baseline 4 is appropriate given the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Voluntarily leave your current family group') with the resource (family group) and actor ('your'). It clearly distinguishes from sibling tools like remove_family_member (which removes others) and create_family_group (creation). The owner exception also sharpens the scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it — when you want to leave your own family group. It provides an exclusion: group owners cannot leave and are pointed to deleting the group instead. However, it does not explicitly name alternative tools (e.g., remove_family_member) or give broader when-to-use context, so it falls short of a perfect 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
link_goal_planGoals: Link goal planAInspect
Link (or unlink) a zoninga goal to a Visions & Goals plan.
Args:
goal_id: The zoninga goal id (owned by the user).
vng_goal_id: The V&G plan id to link — verified against the
user's linked V&G account. Pass null/omit to UNLINK.
Returns:
{"success": True, "goal_id": ..., "vng_goal_id": ...} or an
error dict.
| Name | Required | Description | Default |
|---|---|---|---|
| goal_id | Yes | ||
| vng_goal_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and destructiveHint=false, but the description discloses behavior annotations miss: the same call doubles as an UNLINK operation, and the target is verified against the user's linked V&G account. That unlink capability is meaningful context beyond the structured hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action, then structured Args/Returns blocks. Efficient, though the Args/Returns docstring formatting is slightly verbose for a two-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation with no output schema, the description covers both parameters, the unlink path, the account-verification requirement, and the return shape. Annotations cover the safety profile, so little is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the full burden, and it delivers: goal_id is described as the user-owned goal id, and vng_goal_id as the verified plan id where null/omit triggers unlinking. This adds real meaning the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: linking/unlinking a 'zoninga goal' to a 'Visions & Goals plan'. This is clearly distinct from siblings like create_goal_plan, get_goal_plan, and list_goal_plans, which an agent can tell apart without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives one concrete usage rule (pass null/omit to UNLINK), which is helpful, but offers no when-to-use guidance against alternatives or prerequisites for when linking is appropriate. Usage context is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_accountsAccounts: List accountsARead-onlyIdempotentInspect
List all financial accounts for the user.
Args:
account_type: Filter by type (checking, savings, credit_card, loan, investment, other)
include_hidden: Include hidden accounts (default: False)
limit: Maximum number of accounts to return (default: 50, max: 200).
offset: Number of accounts to skip for pagination (default: 0).
Returns:
Dictionary with a paginated accounts list (``total_count`` /
``has_more`` / ``offset`` / ``limit``) and a ``summary`` that
reflects ALL active accounts — household-wide, regardless of the
page AND of ``account_type`` (its net_worth is the real net
worth in every case). With ``account_type`` set, ``filtered_subtotal`` adds the
count and total balance of the matching active accounts. Debt
accounts include Plaid enrichment: is_overdue, balance_status,
loan_status, next_payment_due_date.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| account_type | No | ||
| include_hidden | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With read-only annotations already covering safety, the description adds substantial behavioral context: pagination fields (total_count, has_more, offset, limit), a summary that is household-wide regardless of page or account_type, filtered_subtotal behavior, and debt-account Plaid enrichment fields. This goes well beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a one-line purpose, then structured Args and Returns sections. Every sentence is relevant given the lack of an output schema and the need to document parameters explicitly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-parameter read-only list tool with no output schema, the description is complete: it explains parameters, return structure, pagination, summary semantics, and debt enrichment. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden and does so completely. It documents account_type allowed values, include_hidden default, limit default and max, and offset default, adding meaning not present in the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'List all financial accounts for the user.' It is clear what the tool does, but it does not explicitly distinguish itself from account-related siblings such as list_shared_accounts or get_account_details.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the statement 'List all financial accounts for the user,' and the Args section clarifies filtering options. However, there is no explicit guidance on when to use this tool versus alternatives like list_shared_accounts or get_account_details.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_agent_chat_sessionsAgent Chat: List agent chat sessionsARead-onlyIdempotentInspect
List the user's past AI assistant chat sessions.
Useful for answering "what did we discuss last time?" or helping
the user find a specific past conversation. Returns the most
recently updated sessions first. Transcripts are not included.
Args:
limit: Maximum number of sessions to return (1-50, default 10).
Returns:
{"sessions": [{id, title, is_active, message_count, ...}, ...],
"count": N}
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world). The description adds meaningful behavior beyond that: most-recently-updated-first ordering, exclusion of transcripts, and the return payload shape.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose, then usage context, then structured Args/Returns sections. No filler sentences; every line adds information an agent can act on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with one optional parameter and no output schema, the description supplies ordering, content exclusions (no transcripts), parameter bounds, and the return structure — everything needed to invoke and interpret it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden, and it does: it documents the sole parameter's meaning, range (1-50), and default (10), fully compensating for the undocumented schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the user's past AI assistant chat sessions') and clarifies scope by noting transcripts are not included, which implicitly separates it from get_agent_chat_session. It doesn't explicitly name that sibling, so it stops short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage contexts ('what did we discuss last time?', finding a specific past conversation), which tells the agent when this tool is appropriate. There is no explicit when-not guidance or named alternative, but the trigger conditions are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_badgesGamification: List badgesARead-onlyIdempotentInspect
List all badges with their earned/unearned status for the user.
Args:
category: Filter by badge category (milestone, goal, budget,
financial_health, family, streak, seasonal)
Returns:
All badges grouped by category with earned status and dates
| Name | Required | Description | Default |
|---|---|---|---|
| category | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds expected return-shape context (grouped by category, earned status, dates) but no new behavioral warnings. There is no contradiction between the description and annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: a clear one-line purpose, a short Args section, and a brief Returns section. It avoids filler and front-loads the most important information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with no output schema, the description adequately covers what is returned: badges grouped by category with earned status and dates. Minor gaps include no explicit note that category is optional or details about the exact response shape, but the schema default and low complexity keep this acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides only a generic 'category' property with no description, so the description carries the full semantic burden. It explains that category filters badges and enumerates the allowed values: milestone, goal, budget, financial_health, family, streak, and seasonal.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'List all badges' for the user, with earned/unearned status and grouping by category. This clearly differentiates it from similar gamification tools like get_gamification_summary, get_points_and_level, and list_milestones.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: call this when the agent needs the user's complete badge inventory with earned status, optionally filtered by category. However, it does not explicitly name alternatives or state when not to use it, leaving the selection logic mostly to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_budgetsBudgets: List budgetsARead-onlyIdempotentInspect
List all budgets for the user.
Args:
month: Filter by month in YYYY-MM format (defaults to current month)
limit: Maximum number of budgets to return (default: 50, max: 200).
offset: Number of budgets to skip for pagination (default: 0).
Returns:
List of budgets with spending progress, plus ``total_count`` /
``has_more`` / ``offset`` / ``limit`` pagination metadata. The
``summary`` totals cover ALL budgets for the month, not just the
returned page.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| month | No | ||
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive and closed-world, so safety is covered. The description adds genuinely non-obvious behavioral context: that the ``summary`` totals cover ALL budgets for the month rather than just the returned page, plus the pagination metadata shape (total_count/has_more/offset/limit).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then organized into Args/Returns sections. Efficient overall, though the Args block restates defaults already present in the schema, adding mild redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly fills the gap by describing the return payload and pagination metadata. All three parameters are documented, and the total-vs-page caveat covers the main ambiguity an agent would face.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does: it documents all three parameters with the month format (YYYY-MM), the default-to-current-month behavior, the limit default of 50 with a max of 200, and the offset semantics. It largely duplicates the schema defaults but supplies missing format and bound constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb and resource: "List all budgets for the user." An agent knows exactly what it retrieves. However, it does not differentiate itself from related siblings such as get_budget_history, get_budget_streaks, or create_budget/update_budget/delete_budget.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description never says when to use this tool versus alternatives like get_budget_history or get_budget_streaks, nor does it state prerequisites. It only describes the retrieval mechanics, leaving usage to be inferred from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_categoriesCategories: List categoriesARead-onlyIdempotentInspect
List all available categories.
Args:
category_type: Filter by type (income or expense)
Returns:
List of categories grouped by type
| Name | Required | Description | Default |
|---|---|---|---|
| category_type | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the 'grouped by type' return behavior, which is useful, but doesn't add substantial behavioral context beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: a one-line purpose, an Args section, and a Returns section. Every sentence earns its place, and the core purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with one optional parameter and no output schema, the description covers the purpose, the filter, and the return shape. It could specify more detail about the category object structure, but it is largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the burden for explaining the parameter. It explains that category_type is a filter and suggests valid values ('income or expense'), which is meaningful beyond the bare schema definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the operation: 'List all available categories' – a specific verb and resource. It is unambiguous about scope, though it doesn't explicitly differentiate from sibling tools like list_tags or list_forum_categories beyond the resource name itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives the optional filter context ('Filter by type (income or expense)') but provides no explicit when-to-use guidance or exclusions relative to alternatives. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_challengesGamification: List challengesARead-onlyIdempotentInspect
List challenges: active (with progress), available to join, and completed.
Returns:
Active challenges with progress, available challenges, and completed count
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the operation read-only, idempotent, and non-destructive. The description adds useful behavior beyond those facts by specifying what the listing contains: active challenges with progress, challenges available to join, and a completed count. No contradiction with annotations exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short, scannable, and front-loads the core meaning. The Returns block largely repeats the first sentence, creating minor redundancy, but the overall structure is still efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only list operation with no output schema, the description gives enough of the return shape to set expectations. It is complete enough for simple invocation, though it doesn't specify ordering or formatting.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, so there is no parameter semantics burden for the description to carry. The baseline of 4 is appropriate since nothing is left undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List challenges') and goes on to define the three result categories: active with progress, available to join, and completed. This clearly distinguishes it from sibling mutation tools like join_challenge, leave_challenge, and create_spending_challenge.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance is given for when to use this tool versus alternatives such as get_gamification_summary, get_leaderboard, or list_badges. The read-only listing intent is implied, but there are no exclusions or alternative-routing statements.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_debt_payment_rulesDebt Payments: List debt-payment rulesARead-onlyIdempotentInspect
List the user's debt payment rules.
Args:
is_active: Filter by active/inactive (default: show all)
Returns:
{rules: [...], count: N}
| Name | Required | Description | Default |
|---|---|---|---|
| is_active | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only, idempotent, non-destructive behavior. The description adds the return contract ({rules: [...], count: N}) and the default 'show all' behavior for is_active, which is useful beyond the annotations. It omits ordering/pagination but that is minor for this simple list.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The action is stated in the first sentence, and the Args/Returns sections are compact and scannable. There is no filler or repeated schema boilerplate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-optional-parameter read-only list tool with annotations, the description covers the action, filter semantics, default behavior, and return shape. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden. It explains that is_active filters by active/inactive and that the default is to show all, giving the parameter meaningful behavior rather than just a nullable boolean.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific operation and resource: it lists the user's debt payment rules, which distinguishes it from rule mutation tools (create/update/delete_debt_payment_rule) and from list_debt_payments. The scope ('the user's') is explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The listing intent is clear from the verb and resource, but the description gives no explicit when-to-use or when-not-to-use guidance and names no alternatives. An agent would infer use from the name rather than being told.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_debt_paymentsDebt Payments: List debt paymentsBRead-onlyIdempotentInspect
List matched debt payments (pending / confirmed / rejected).
Args:
status: Filter by status (pending, confirmed, rejected). Default: all.
limit: Max number of records to return (default 20, max 200)
Returns:
{payments: [...], count: N}
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| status | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds the default limit and the return shape ({payments: [...], count: N}), which is useful, but it doesn't disclose pagination behavior beyond the limit parameter or clarify what 'matched' means in terms of matching state. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core purpose, followed by parameter details and return shape. Every sentence earns its place, though the Args/Returns formatting is slightly verbose for a simple two-parameter tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with two optional parameters and no output schema, the description covers the essentials: what it lists, the filters, and the return shape. It doesn't explain the meaning of 'matched' debt payments or how this relates to the debt payment matching workflow, which could matter for an agent deciding between this and confirm_debt_payment_match.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter meaning. It does explain status values (pending, confirmed, rejected) and limit defaults (20, max 200), which adds value beyond the bare schema. However, it doesn't specify the exact accepted status string format beyond the examples, and the schema itself has no enums, so there is some ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('List matched debt payments') and enumerates the status filter values, which distinguishes it from related tools like list_debt_payment_rules and confirm_debt_payment_match. It doesn't explicitly name a sibling to differentiate from, but the scope is clear enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by listing the status filter and limit, but it doesn't explicitly state when to use this tool versus alternatives like list_payment_history or confirm_debt_payment_match. There is no when-not-to-use guidance, so an agent must infer the context from the name and sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_family_chat_messagesFamily: List family chat messagesARead-onlyIdempotentInspect
Get recent family chat messages.
Args:
after_id: Only return messages after this ID (for polling)
limit: Maximum messages to return (default: 50, max: 100)
Returns:
List of chat messages
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| after_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idiompotent/destructive safety. The description adds useful behavior beyond annotations: after_id semantics for polling and the limit max of 100, which is not in the schema. It does not mention ordering or pagination beyond these, but the 'recent' keyword implies a common ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A focused docstring with a one-line purpose, concise parameter explanations, and a returns line. No fluff, and the key purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list tool with two optional parameters and no output schema, the description covers purpose, parameters, and return type. It could be more explicit about ordering (e.g., newest first), but 'recent' and the polling hint cover most expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries full burden. It fully explains both parameters: after_id for incremental polling and limit with default and max. This adds meaning the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'Get recent family chat messages.' It is distinct from send/edit/delete family chat message siblings, though it does not explicitly differentiate itself. The resource is specific enough that an agent can identify it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage through 'for polling' and the after_id parameter, but it does not explicitly state when to use this tool versus alternatives like send_family_chat_message or get_family_activity. No exclusions or alternative routing are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_family_invitationsFamily: List family invitationsARead-onlyIdempotentInspect
List pending family invitations sent to you or by you.
Returns:
Invitations you've received and invitations you've sent.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint, idempotentHint, and non-destructive behavior. The description adds useful context by specifying 'pending' and clarifying that both received and sent invitations are returned, but it does not go beyond that into other behavioral details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief, front-loaded with the main purpose, and then clarifies the return content. Every sentence contributes meaning without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only listing tool with strong annotations, the description is nearly complete: it states what is listed and what will be returned. A small gap is that it does not describe the expected shape or fields of an invitation, but this is not critical for selecting or invoking the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the empty schema fully covers invocation needs. The description doesn't need to explain parameters, and the no-input behavior is clear from the schema context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'List pending family invitations', and further scopes it to 'sent to you or by you'. This clearly distinguishes it from action-oriented siblings like accept, decline, send, and revoke invitation tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the tool's basic use context obvious: it lists pending invitations in both directions. However, it does not explicitly state when to choose this over related tools such as accept_family_invitation, decline_family_invitation, or send_family_invitation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_family_membersFamily: List family membersBRead-onlyIdempotentInspect
List all members in your family group with roles and join dates.
Returns:
List of family members with role and status info.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, covering the safety profile. The description adds the return content (roles, join dates, status) which is helpful but does not disclose other behavioral traits such as whether the list is ordered, paginated, or scoped to the current user's family group. It adds some value beyond annotations but not substantial context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences and front-loaded with the action and scope. The 'Returns' line clarifies the output, though it partially duplicates the first sentence ('roles and join dates' vs 'role and status info') and introduces a minor inconsistency. Overall it is compact and useful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (0 params, no output schema) and annotations covering safety, the description is largely complete. It states what is returned. It could mention prerequisites like family group membership or the absence of filtering, but those are minor given the straightforward nature of the operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters and the schema is empty, so per the rubric the baseline is 4. The description doesn't need to explain any parameters. It adds no parameter-specific meaning, but none is required.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'List all members in your family group with roles and join dates.' This clearly identifies the tool's function. It does not explicitly differentiate from siblings like get_family_group or list_family_invitations, so it falls short of the top score for sibling distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention prerequisites (e.g., must belong to a family group) or exclude cases like viewing invitations or group details. The only implicit hint is the tool name and title, which is insufficient for a clear usage policy.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_forum_categoriesForum: List forum categoriesARead-onlyIdempotentInspect
List active forum categories.
Returns:
``{"categories": [{slug, name, description, icon, is_staff_only}],
"count": N}``. A ``slug`` identifies the category when listing
or creating posts.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the safety profile is covered. The description adds real behavioral value by disclosing the response shape and the fact that only 'active' categories are returned, plus what a slug is for.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The one-line purpose is front-loaded and the returns block is compact. The indented docstring formatting is slightly noisy but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description responsibly supplies the return structure (categories, count, field names). It omits ordering/pagination, but for a zero-parameter, non-paginated list this is essentially complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the baseline is 4. The description adds no parameter detail because none exists, and none is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List active forum categories'), which clearly separates it from list_forum_posts and create_forum_post. It does not explicitly name a sibling, but the resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The note that 'a slug identifies the category when listing or creating posts' implies the downstream use of the result, which is helpful context. However, it never states when to call this versus alternatives or any prerequisites, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_forum_postsForum: List forum postsARead-onlyIdempotentInspect
List recent forum posts, newest first (pinned posts first).
Args:
category_slug: Filter by category slug (optional). An unknown slug
raises a validation error.
limit: Max posts to return (default 20, max 100)
offset: Number of posts to skip for pagination (default: 0)
Returns:
``{"posts": [{id, category_slug, author_id, title, content,
is_pinned, is_locked, reply_count, created_at, updated_at}],
"total_count": N, "has_more": bool, "offset": N, "limit": N}``.
Staff-only categories are excluded for non-staff.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| category_slug | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive), but the description adds real behavioral context: pinned-first ordering, staff-only category exclusion for non-staff, and that an unknown slug raises a validation error.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then Args and Returns sections. Slightly long due to enumerating return fields, but every part is useful and well-organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the full return shape (posts array fields, total_count, has_more, offset, limit) and pagination semantics. Annotations cover safety, so nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the burden and it does: it explains category_slug filtering plus the validation-error behavior, limit default 20/max 100, and offset as pagination skip count. All three params gain meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List recent forum posts') plus ordering scope (newest first, pinned first). It is clearly distinguishable from get_forum_post and create_forum_post by verb, but it never names a sibling or explicitly contrasts itself with list_forum_categories.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: the tool browses posts and can be narrowed with an optional category filter. There is no explicit when-to-use vs alternatives guidance (e.g., when to reach for get_forum_post instead, or how it relates to list_forum_categories).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_goal_plansGoals: List goal plansARead-onlyIdempotentInspect
List the user's goal PLANS on Visions & Goals (the sister site for
clarifying and planning goals — SMART criteria, action steps).
These are qualitative plans, separate from the money-tracking
zoninga goals. A plan and a goal can be linked via ``vng_goal_id``.
Returns:
{"goals": [{id, title, timeframe, priority, target_date,
is_achieved, smart_score, action_plan {count,
completion_percentage}, edit_url, ...}], "count": N} — or an
error with a hint when the accounts aren't linked yet.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, so the safety profile is covered. The description adds genuinely new behavioral context: it returns a structured payload, and notably returns an error with a hint when the accounts aren't linked yet — an edge case not captured by annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and scoping are front-loaded in the first sentence, followed by the disambiguation and the return shape. The Returns block is verbose but useful given there is no output schema; the definition is slightly padded but each part earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by sketching the return object (fields, count, action_plan sub-object) and flagging the unlinked-accounts error case. Combined with the annotations covering the safety profile, an agent has enough to call and interpret it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the baseline a 4 is appropriate. There are no inputs to explain, and the description correctly focuses on output shape instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (goal PLANS), and goes further by scoping them to the Visions & Goals sister site. It explicitly distinguishes these qualitative plans from the money-tracking goals, which an agent needs to tell list_goal_plans apart from list_goals and get_goal_plan.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear situational context: these are qualitative planning artifacts (SMART criteria, action steps) tied to Visions & Goals, not the financial goal objects handled by list_goals. It mentions the vng_goal_id linkage, which signals relationships to goal tools. However, it stops short of explicitly stating when to prefer this over list_goals/list_shared_goals.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_goalsGoals: List goalsARead-onlyIdempotentInspect
List the user's financial goals.
Default scope: active + completed goals (cancelled goals are
hidden by default to avoid cluttering responses with abandoned
targets, which accumulate over time). Pass
``include_cancelled=True`` or an explicit ``status="cancelled"``
to see them.
Args:
status: Filter by a single status — "active", "completed",
or "cancelled". When provided, ``include_cancelled``
is ignored.
include_cancelled: If True (and ``status`` is None), include
cancelled goals in the default active+completed listing.
limit: Maximum number of goals to return (default: 50, max: 200).
offset: Number of goals to skip for pagination (default: 0).
Returns:
List of goals with progress info, plus ``total_count`` /
``has_more`` / ``offset`` / ``limit`` pagination metadata.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| status | No | ||
| include_cancelled | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already establish read-only, idempotent, and non-destructive behavior, and the description adds substantial behavioral context: cancelled goals are hidden by default to avoid clutter, include_cancelled and status interact with a defined precedence, and pagination metadata is returned. This clearly exceeds what annotations alone provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then provides dense, useful parameter and return-value documentation. Every section earns its place, and there is no redundant repetition of the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description appropriately explains what the tool returns: a list of goals with progress info plus pagination metadata. It also covers default scoping, cancellation handling, and all parameter semantics. For a read-only listing tool, this is complete enough for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden of explaining the parameters. It compensates thoroughly by defining status values ('active', 'completed', 'cancelled'), explaining the include_cancelled/status precedence, and adding limit default/max and offset pagination details that are absent from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action and resource: 'List the user's financial goals.' It also adds useful scoping detail about active/completed/cancelled goals. However, it does not explicitly distinguish itself from sibling tools like list_shared_goals or list_goal_plans, so it falls short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear guidance on when to use status vs include_cancelled and explains the default behavior, so an agent knows how to adjust the call. It does not explicitly discuss when to prefer this tool over related list tools, but the context is clear enough for most goal-listing scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_hard_assetsEstate: List hard assetsARead-onlyIdempotentInspect
List all hard assets (real estate, vehicles, valuables) for net worth tracking.
Returns:
List of hard assets with categories, values, and depreciation info.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish read-only/idempotent/non-destructive behavior. The description adds useful output context ('categories, values, and depreciation info') and claims completeness ('all hard assets'), but does not mention pagination, ordering, or auth semantics. This is adequate for a simple read operation but no more.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Short, front-loaded, and free of filler. The 'Returns:' line is purposeful because there is no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless, read-only list with rich annotations, the description covers the purpose, scope, and return shape. Nothing essential is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the schema carries no burden and the description's 'all hard assets' reinforces the unfiltered scope. Baseline for zero-parameter tools applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') with a clear resource ('all hard assets') and enumerates covered subtypes ('real estate, vehicles, valuables'), which distinguishes it from create/update/delete_hard_asset and from aggregate net-worth tools. The 'for net worth tracking' purpose adds further context without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It says the list serves 'net worth tracking,' giving a general context for use, but it does not explain when to prefer this over related getters like get_net_worth or get_balance_sheet, nor does it state exclusions. Usage is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_milestonesMilestones: List milestonesARead-onlyIdempotentInspect
List all milestones for a goal, ordered by target amount.
Args:
goal_id: The goal whose milestones to list
Returns:
Ordered list of milestones with target amount and reached status
| Name | Required | Description | Default |
|---|---|---|---|
| goal_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and idempotentHint=true, so the safety profile is covered. The description adds useful behavioral detail beyond annotations: milestones are returned ordered by target amount and include both target amount and reached status. This is meaningful for an agent predicting the tool's output and behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three short, purposeful sections with no filler. The main verb-resource statement is front-loaded, followed by a minimal parameter description and a concise return summary. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only, non-destructive list operation with no output schema, the description is complete: it states the operation scope, required input, ordering, and returned fields. No pagination, error, or mutation details are necessary. The absence of an output schema is mitigated by the explicit return description.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema only provides an integer type and title 'Goal Id', with 0% schema description coverage. The description compensates by explaining that goal_id is 'The goal whose milestones to list,' clarifying the parent-entity relationship. For a single, simple integer parameter, this is sufficient semantic guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'List all milestones for a goal.' It clearly distinguishes this read-only listing tool from siblings like create_milestone, update_milestone, delete_milestone, list_goals, and list_goal_plans. The scope ('for a goal') and ordering detail ('by target amount') make the tool's purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The intended use is implied by the phrase 'for a goal' and the required goal_id parameter, but no explicit when-to-use guidance or alternatives are mentioned. It does not exclude other list tools or explain when this tool is preferable to related milestone or goal retrieval tools. The guidance is serviceable but relies on inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_notificationsNotifications: List notificationsARead-onlyIdempotentInspect
List notifications for the user.
Args:
unread_only: Only show unread notifications (default: True)
limit: Maximum number of notifications to return (default: 20, max: 100)
offset: Number of notifications to skip for pagination (default: 0)
Returns:
List of notifications plus ``unread_count`` (the full unread
total, regardless of filter) and ``total_count`` / ``has_more`` / ``offset`` / ``limit``
pagination metadata for the current filter.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| unread_only | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive behavior, so the safety profile is covered. The description adds genuinely useful context beyond that: unread_count is the full unread total regardless of the filter applied, and pagination metadata is returned for the current filter. That subtlety would be easy to get wrong otherwise.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Docstring layout with Args and Returns is well organized and front-loaded with the one-line purpose. Every line carries information, though the Args/Returns scaffolding is slightly more verbose than a prose equivalent would need to be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description responsibly documents the return payload including unread_count, total_count, has_more, offset, and limit. The main remaining gap is that it does not describe what fields an individual notification object contains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the burden and largely does: it explains unread_only's meaning and default, limit's default and max (100), and offset's pagination role. Only the absence of any note about ordering or accepted ranges below zero keeps it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (notifications) scoped to the user, which is unambiguous. It does not explicitly differentiate itself from siblings like get_unread_notification_count or mark_notification_read, but the verb+resource pairing is clear enough to place it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the unread_only default and pagination params, and the reader can infer this is the general listing tool. However, it never says when to use this versus get_unread_notification_count or when filtering is unnecessary, and no alternatives are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_payment_historySubscription: List payment historyARead-onlyIdempotentInspect
List the user's recent Stripe payment history.
Args:
limit: Maximum rows (default: 10, max: 100)
Returns:
``{"payments": [...]}`` with id, status, amount, currency,
created_at, plan_name, description, and invoice_pdf_url.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds the return shape (list of payments with fields) and the 'recent' scoping, which is useful context beyond annotations. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief and front-loaded with the core purpose. The Args and Returns sections add necessary detail without redundancy. Every sentence earns its place, and the structure is clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with one parameter and no output schema, the description provides the return format and parameter constraint. It does not mention pagination or any filtering beyond 'recent', but given the simplicity, it is adequately complete. The lack of an output schema is mitigated by the explicit Return section.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one parameter (limit) with no description, so schema coverage is 0%. The description compensates fully by explaining 'Maximum rows (default: 10, max: 100)', giving both the meaning and constraints. This is complete and unambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'List the user's recent Stripe payment history.' It is specific about the resource (Stripe payments) and distinguishes it from siblings like list_payment_links (payment links) and get_subscription_status (current status). No ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention any sibling tools, prerequisites, or conditions. The agent must infer from the name and description when to choose this over other list/get tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_payment_linksAgent: List payment linksARead-onlyIdempotentInspect
List the user's saved payment links — external URLs the user provided
for specific payment contexts (e.g. "child support" → the state
child-support portal).
For a matching payment context, a saved payment link takes precedence
over the source account's bank login link (institution_url).
``url_host`` is the hostname, suitable as visible link text so the
destination is shown.
Returns:
{"payment_links": [{id, label, url, url_host, updated_at}, ...],
"count": N}
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint and destructiveHint=false, so safety is covered. The description adds genuine domain behavior beyond that: the precedence rule over institution_url and the meaning of url_host as suitable display text. It stops short of describing pagination or ordering, but adds real context against a lowered bar.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then adds precedence semantics and a returns block. Every sentence carries information, though the Returns block is a touch verbose for a zero-parameter list tool. No filler or tautology.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Low-complexity, read-only, zero-parameter tool with no output schema, and the description compensates by documenting the return object shape (payment_links array with id, label, url, url_host, updated_at). Nothing essential for correct invocation is missing, though pagination/ordering behavior is left unspecified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the schema baseline of 4 applies. There are no parameters to misinterpret, and the description usefully documents the shape and fields of the returned objects instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('payment links'), and goes further by defining what a payment link actually is (an external URL the user provided for a specific payment context). The intent is unambiguous against siblings like save_payment_link and delete_payment_link, though those siblings are not named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated. The note that a saved payment link takes precedence over institution_url for a matching context explains the data's semantics but does not tell the agent when to call this tool versus list_accounts or get_account_details. No explicit when/when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_plaid_itemsPlaid: List linked banksARead-onlyIdempotentInspect
List all Plaid bank connections (items) for the user.
Each item represents one institution (one bank login). A single
item can hold multiple Account rows (checking + savings + credit
card from the same bank).
Args:
include_inactive: Include unlinked / soft-deleted items (default: False)
Returns:
``{"items": [...], "total_count": N}`` where each item carries
id, institution_name, products (consented), last_synced_at,
is_active, error_code/message, consent_expiration_time, and
the count of currently-linked accounts.
| Name | Required | Description | Default |
|---|---|---|---|
| include_inactive | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds meaningful behavioral context: it explains what unlinked/soft-deleted means via include_inactive, and it details the return payload fields, including consent_expiration_time and error_code/message. This goes beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a one-sentence summary, a clarifying paragraph, and explicit Args/Returns sections. It front-loads the purpose and adds value without fluff. Every sentence contributes to correct invocation and interpretation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter read-only tool, the description is complete: it names the parameter, its default, and the return shape with all key fields. No output schema exists, so the Returns section appropriately covers what an agent would need to interpret results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverageebbovar, so the description carries full responsibility for explaining the include_inactive parameter. It does so clearly in the Args section: it defines 'unlinked / soft-deleted items' and states the default (False). The single parameter is fully covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and resource ('Plaid bank connections (items)'), with a clear scope statement ('for the user'). It also details what each item representsholistically, distinguishing it from account-level tools like list_accounts and the sibling get_plaid_item.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states the operation is user-scoped and explains the optional include_inactive flag, but it does not explicitly name alternative tools or exclusion criteria. Since the context is unambiguous and the sibling set includes get_plaid_item, some routing guidance could be inferred, but it is not explicitly provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_recurring_transactionsRecurring Transactions: List recurring templatesARead-onlyIdempotentInspect
List the user's recurring transaction templates.
Templates are the parent rows that the daily cron spawns child
transactions from. Spawned children appear in the regular
transaction listing, not here.
Returns:
``{"templates": [...], "total_count": N}`` — each template
includes id, description, amount, frequency, start/end date,
and the running ``instances_generated`` count.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds genuine conceptual context beyond that: templates are parent rows spawned by a daily cron, and the instances_generated counter reflects running output. It does not mention pagination or sorting, which is a minor gap for a list endpoint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences plus a compact Returns block. The core purpose is front-loaded in the first sentence, and the remaining content is the disambiguation and return shape, each earning its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and no output schema, the description carries the whole burden and does so: it names the return envelope (templates, total_count) and the per-template fields, and explains the parent/child relationship. Nothing essential for calling this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate at the parameter level; baseline for a 0-param tool is 4. No misleading parameter claims are made.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the user's recurring transaction templates') and immediately disambiguates from the far larger sibling list_transactions by explaining that spawned children appear there, not here. An agent can distinguish this from list_transactions without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly indicates the context in which this tool is correct (retrieving parent templates) and where the spawned children live instead, which routes the agent away from the wrong listing. It stops short of explicit when-not-to-use guidance or naming alternative tools (e.g. list_transactions, get_subscription_status) directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_remindersAgent: List scheduled remindersARead-onlyIdempotentInspect
List the user's scheduled reminders — texts they asked the agent to
resurface on a schedule (delivered in the morning brief, and in the
brief email when they've opted in).
Args:
include_inactive: Also include delivered one-time and deactivated
reminders (default False).
Returns:
{"reminders": [...], "delivery_note"?: str} — delivery_note appears
when the brief cadence means reminders won't actually be delivered.
| Name | Required | Description | Default |
|---|---|---|---|
| include_inactive | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive behavior, but the description adds real behavioral context: reminders are resurfaced in the morning brief and optional brief email, inactive reminders are excluded by default, and delivery_note appears when the brief cadence prevents delivery. This goes beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the core action, and every sentence earns its place: the definition, the parameter, and a return caveat. No filler or repetition of schema data.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only list tool, this is complete: it defines what is listed, how reminders are delivered, what the parameter controls, and what the response contains, including the edge case where delivery_note appears. No output schema exists, so the return documentation is necessary and sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 0%, the description fully documents the only parameter, include_inactive, explaining that it includes delivered one-time and deactivated reminders and that it defaults to False. This adds meaning well beyond the schema's bare boolean and default value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource: 'List the user's scheduled reminders', and further clarifies them as 'texts they asked the agent to resurface on a schedule'. It also identifies their delivery channels, making the tool's role unique among siblings like list_notifications.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives or when not to use it. The description implies the context by defining scheduled reminders, but it never names related tools such as create_reminder or delete_reminder or states conditions that should route an agent elsewhere.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_scenariosScenarios: List what-if scenariosARead-onlyIdempotentInspect
List the user's saved what-if scenarios (active, not deleted).
Each scenario is a "potential choice" that auto-expires 60 days after
creation. The full projected impact is available per scenario by id.
Returns:
{"scenarios": [{id, name, status, item_count, expires_at, ...}],
"count": N}
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior, so the safety profile is covered. The description adds genuinely useful context beyond that: scenarios auto-expire 60 days after creation, deleted ones are excluded, and the return payload shape is shown.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then a compact note on expiry, then the return shape. The inline expandable detail is acceptable but the Returns block adds a little bulk that could be trimmed given there is no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and no output schema, the description carries the return-value burden and does so by sketching the response shape ({scenarios: [...], count}). Combined with the expiry and soft-delete notes, an agent has everything needed to call and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. There is nothing in the schema to expand on, and the description does not need to compensate for missing parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (user's saved what-if scenarios) and narrows scope to active, not-deleted entries. This cleanly distinguishes it from siblings like get_scenario, create_scenario, and delete_scenario.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description conveys when to use it (enumerate saved scenarios) and hints at the alternative route by noting 'The full projected impact is available per scenario by id,' which points toward get_scenario. It never names that sibling explicitly or states exclusions, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_subscription_plansSubscription: List subscription plansARead-onlyIdempotentInspect
List all active subscription plans (Free, Premium monthly/yearly).
Returns:
``{"plans": [...]}`` — each plan includes id, name, plan_type,
price, currency, billing_interval, and the plan's quotas / feature
flags, which together answer "what do I get if I upgrade?" Monthly and yearly Premium are
separate plan rows (distinguished by ``billing_interval``).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and non-open-world, so safety is covered. The description adds behavior the annotations cannot: the exact return shape ({'plans': [...]}), the fields each row carries, and the non-obvious fact that monthly and yearly Premium are separate rows keyed by billing_interval.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action, followed by a compact return description; no filler sentences. The 'Returns:' block is slightly bulky relative to a no-param list tool, but every line conveys new information about the payload.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no input parameters and no output schema, the description is the sole source of payload information, and it supplies both the top-level envelope and the per-plan field set. Nothing an agent needs to call and interpret this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Zero parameters, so there is nothing for the description to disambiguate; the documented return fields (id, name, plan_type, price, currency, billing_interval) are value-add rather than parameter compensation. Baseline 4 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List all active subscription plans') and further pins down scope by naming the concrete plan rows (Free, Premium monthly/yearly). It distinguishes itself from neighbors like get_subscription_status and list_payment_history because it is the catalog of available plans rather than the user's own billing state.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the framing question 'what do I get if I upgrade?', which tells the agent this is a pre-purchase comparison tool. However, no alternative is named (e.g., get_subscription_status for current plan, start_checkout for upgrading), so the agent must infer routing from sibling names alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_support_ticketsSupport: List support ticketsARead-onlyIdempotentInspect
List the user's own support tickets, most recently active first.
API docs: https://zoninga.com/ai/#support-tickets
Closed tickets are hidden by default (they're soft-archived once
closed). Pass a ``status`` to change that:
Args:
status: ``None`` (default) hides closed tickets. ``"all"`` shows
every ticket including closed. A specific status —
``open``, ``in_progress``, ``resolved``, ``closed`` — filters
to exactly that. Any other value raises a validation error.
limit: Max tickets to return (1-100, default 20).
Returns:
``{"tickets": [{id, subject, category, category_display, status,
status_display, priority, priority_display, message_count,
created_at, updated_at}], "count": N}``.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| status | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/non-destructive, so the bar is lower, yet the description still adds real behavior: closed tickets are soft-archived and filtered out by default, results are sorted by most recent activity, and invalid status values raise a validation error. It does not mention pagination or total-vs-returned count semantics, which is the remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and default behavior are front-loaded, and the Args/Returns structure keeps it scannable. The docstring formatting and API-docs URL add slight bulk, but every sentence conveys usable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the exact return shape (tickets array with field list plus count), covering what an agent needs to consume the result. Combined with the filtering and limit semantics, nothing essential is missing for a 2-param read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden and does so: it enumerates every accepted status value (None, 'all', open, in_progress, resolved, closed), states the default, describes the limit range (1-100, default 20), and warns that other values error. This is meaning well beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List the user's own support tickets') plus a scope and sort order ('most recently active first'). This clearly distinguishes it from get_support_ticket, create_support_ticket, and reply_to_support_ticket without needing their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explains the default behavior (closed tickets hidden) and the condition that changes it via an explicit status list, which is solid context for when to use which filter. It stops short of naming sibling tools or saying when to prefer get_support_ticket for a single record, but the usage context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_tagsTags: List tagsARead-onlyIdempotentInspect
List all tags the user has defined.
Returns:
``{"tags": [...], "total_count": N}`` — each tag has id,
name, color (hex), and how many transactions currently carry it.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is established. The description adds valuable behavioral context by specifying the return format ({"tags": [...], "total_count": N}) and the fields of each tag (id, name, color hex, transaction count), which goes beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences: the first states the purpose, the second details the return format. It's front-loaded, with no extraneous words. Every sentence earns its place, making it highly concise and efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter list operation, the description fully covers what an agent needs: what the tool does and what the output looks like, including field-level detail. There's no output schema, so the description correctly compensates by explaining the return structure. Annotations cover safety, so nothing is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so there is nothing to document. Baseline 4 applies because no parameter semantics are needed; the description correctly omits any param details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists 'all tags the user has defined', specifying the verb (list), the resource (tags), and scope (user-defined). It distinguishes itself from sibling mutation tools like create_tag/delete_tag/update_tag by being a read operation, so there's no ambiguity about what it does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The context is clear: use this to retrieve all tags. Since there is no alternative listing tool for tags among siblings, explicit when-not or alternative guidance isn't necessary; the description's single-purpose nature makes usage obvious. It lacks explicit exclusions but doesn't need them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_themesGamification: List themesARead-onlyIdempotentInspect
List all available color themes with the user's active theme.
Returns:
List of themes with active status
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context by specifying the return value: a list of themes with active status.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The core behavior is front-loaded, and the return value is stated succinctly in a separate 'Returns' line.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-argument, read-only listing tool, the description fully covers what is returned and the scope of available themes. Given the annotations, there is no critical missing information needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema carries no burden. Baseline 4 applies because there are no parameter semantics to clarify; the description appropriately focuses on output instead.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('List') and resource ('all available color themes'), and clarifies scope by including the user's active theme and active status. It clearly distinguishes itself from sibling tools like activate_theme and list_badges.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the use case clear: retrieve available themes and see which one is active. It does not explicitly mention when not to use it or point to an alternative like activate_theme, but the purpose is unambiguous for a read-only listing tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_transactionsTransactions: List transactionsARead-onlyIdempotentInspect
List transactions with optional filters.
Args:
account_id: Filter by account ID
category_name: Filter by category name (case-insensitive)
transaction_type: Filter by type (deposit or withdrawal)
start_date: Start date in YYYY-MM-DD format
end_date: End date in YYYY-MM-DD format
search: Search in description
tag_name: Filter by tag name (case-insensitive)
limit: Maximum number of transactions to return (default: 50, max: 200)
offset: Number of transactions to skip for pagination (default: 0)
count_only: If true, return only the total count and totals without transaction details (faster for counting)
Returns:
List of transactions plus ``total_count`` (untruncated) and pagination info.
Each row carries ``economic_class`` (income / expense / transfer /
balance_sheet), the income-vs-spending classification. Raw
``transaction_type`` differs from it on debt accounts: a credit-card
CHARGE is stored as ``deposit`` and a card PAYMENT as ``withdrawal``.
``totals`` (deposits/withdrawals/net) is aggregated over the ENTIRE matching
window in BOTH normal and count_only modes — NOT just the returned page — so a
paginated read reports the full period's income/expenses. The current page's
subtotal is the sum of the per-row ``amount`` values.
SCOPE NOTE: this is a raw transaction listing and INCLUDES rows on
archived (inactive) accounts — like the CSV export, it's the user's
complete history. The cash-flow, spending, uncategorized-summary and
other analytics reports EXCLUDE archived-account transactions, so their
totals differ from a raw list over the same window when the user has
archived accounts. Filtering by an active ``account_id`` matches the
analytics scope.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| offset | No | ||
| search | No | ||
| end_date | No | ||
| tag_name | No | ||
| account_id | No | ||
| count_only | No | ||
| start_date | No | ||
| category_name | No | ||
| transaction_type | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover readOnly/idempotent/non-destructive, yet the description adds substantial behavioral context beyond them: totals are aggregated over the ENTIRE matching window in both normal and count_only modes, economic_class differs from raw transaction_type on debt accounts (charges stored as deposit, payments as withdrawal), and archived accounts are included. These are non-obvious traits an agent cannot infer elsewhere.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose followed by clearly delimited Args/Returns/SCOPE sections, and every block earns its place given the 10 undocumented params and absent output schema. It is somewhat long and the Args list partly restates parameter names, but the added semantics justify the length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description fully specifies the return shape (rows, total_count untruncated, pagination info, economic_class, totals aggregated over the whole window vs page subtotal). Combined with the scope caveat, nothing an agent needs to invoke or interpret this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden and does so well: it documents every parameter's meaning plus format ('YYYY-MM-DD'), case-insensitivity, default 50 / max 200 for limit, pagination semantics for offset, and the speed/count_only tradeoff. This fully compensates for the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb + resource ('List transactions with optional filters') and the SCOPE NOTE explicitly differentiates it from sibling analytics reports (cash-flow, spending, uncategorized-summary) and the CSV export by explaining it includes archived-account rows. An agent can distinguish this raw-listing tool from summary tools without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The SCOPE NOTE tells the agent when this tool's output will diverge from analytics reports and how to reconcile them ('filtering by an active account_id matches the analytics scope'). It gives clear contextual guidance, though it never states an explicit 'use this instead of X' rule, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mark_notification_readNotifications: Mark notification readAInspect
Mark notifications as read.
Args:
notification_id: ID of a single notification to mark as read
mark_all: Set to True to mark all notifications as read
Returns:
Confirmation with count of updated notifications
| Name | Required | Description | Default |
|---|---|---|---|
| mark_all | No | ||
| notification_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations show readOnlyHint=false and destructiveHint=false already, so the mutation aspect is understood. The description adds the return value—'Confirmation with count of updated notifications'—which is useful, but it does not disclose behavior when both arguments are supplied or whether repeated marking is idempotent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the action. The Args and Returns lines are each necessary and directly useful, with no filler or repetition of schema types.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, the description covers the action, both operation modes, and the return shape. The only meaningful gap is the missing rule for what happens when both mark_all and notification_id are provided.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It does: notification_id targets a single notification)Skip, mark_all targets all. This covers the essential semantics, though the interaction between the two parameters is not explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Mark notifications as read.' It is unambiguous and distinct from siblings like delete_notification or get_unread_notification_count, clearly identifying a state-change operation on notifications.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no contextual guidance about when to use this tool versus alternatives. It does not mention when to prefer mark_all over notification_id, nor does it exclude cases like deleting notifications or viewing unread counts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
open_billing_portalSubscription: Open billing portal (returns Stripe URL)AInspect
Create a Stripe Billing Portal session and return the URL.
The portal lets the user update their payment method, view
invoices, and download receipts. Returns ``{"portal_url": "..."}``
or ``{"error": ...}``.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide false hints, so the description carries the burden. It discloses the core side effect (creating a portal session) and the possible return shape (success URL or error), but does not discuss permissions, idempotency, or whether any billing changes happen immediately.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loads the essential action and return type in the first sentence, then gives the portal context and error format. Every sentence carries relevant information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless action with no output schema, the description covers the operation, the user-facing capabilities of the portal, and the return contract. It lacks explicit alternatives or side-effect caveats, but nothing essential is missing for calling it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has zero parameters, which warrants a baseline of 4. The description adds the useful return payload shape ('portal_url' or 'error'), and there are no parameter semantics to clarify.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific action ('Create a Stripe Billing Portal session') and resource, and clearly states it returns a URL. It is unambiguous but does not explicitly differentiate this from related subscription tools like start_checkout or cancel_subscription.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case by explaining the portal lets the user update payment method, view invoices, and download receipts. There is no explicit when-to-use or when-not-to-use guidance, nor any reference to sibling alternatives, so usage guidance is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
override_account_balanceAccounts: Override account balanceAInspect
Manually correct a Plaid-linked account's balance.
When Plaid reports a stale or incorrect balance (mismatched institution
feed, delayed clearing, etc.), the user can override it. The override
is preserved on subsequent Plaid syncs until Plaid reports a genuinely
different (refreshed) balance — at which point the override clears
automatically.
Only works on Plaid-linked accounts. A manual account's balance is
edited directly on the account itself.
Args:
account_id: The Plaid-linked account to override
new_balance: The correct balance (supports negative for overdraft
or credit card balances > credit_limit)
Returns:
Updated account snapshot with new balance and `balance_status`
| Name | Required | Description | Default |
|---|---|---|---|
| account_id | Yes | ||
| new_balance | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-readOnly, non-idempotent, non-destructive, closed-world. The description adds genuinely useful behavioral context the annotations cannot: the override persists across subsequent Plaid syncs and self-clears when Plaid reports a refreshed balance. This lifecycle detail is valuable, though auth/permission requirements and error behavior are unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose followed by the trigger and lifecycle, which is good ordering. The Args/Returns blocks partly restate the schema shape, but since the schema has no descriptions they earn their place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param mutation with no output schema, the description covers purpose, trigger, constraint, lifecycle, and return shape (updated snapshot with balance_status). Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the param burden. It clarifies account_id must reference a Plaid-linked account and that new_balance accepts negatives for overdraft or credit balances exceeding the limit, which meaningfully extends the bare integer/number schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (override) plus resource (Plaid-linked account's balance) and contrasts with the adjacent concept of a manual account whose balance is edited directly. An agent can distinguish this from update_account or accept_plaid_balance from the text alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use it (Plaid reports a stale or incorrect balance from a mismatched feed or delayed clearing) and when not to (manual accounts, which are edited on the account itself). The triggering condition and the excluded alternative are both named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
parse_statementStatements: Parse a PDF statementARead-onlyIdempotentInspect
Parse a base64-encoded PDF bank/credit-card/loan statement and return
the extracted fields WITHOUT saving anything. Premium-only.
This is step 1 of a two-step flow: the extracted values are reviewed
with the user, and a separate apply step writes the confirmed values
to a specific account. Nothing is persisted by this tool.
Args:
pdf_base64: The PDF file encoded as a base64 string (<= 5 MB
decoded). Image-only/scanned PDFs aren't supported.
Returns:
``{"success": True, "fields": {...}, "fields_found": N,
"account_type_hint": "credit_card"|"mortgage"|...|None}`` where
each entry in ``fields`` is
``{"value": "1234.56", "confidence": 0.9, "source_text": "...",
"source": "regex"}`` (the same shape the REST parse returns). The
``value`` strings are the inputs the apply step accepts.
On failure: ``{"error": "..."}`` (not Premium, bad base64, not a
PDF, too large, or unparseable).
| Name | Required | Description | Default |
|---|---|---|---|
| pdf_base64 | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, but the description adds substantial context beyond them: nothing is persisted, the tool is Premium-gated, there is a 5 MB decoded size limit, scanned PDFs are unsupported, and specific failure modes are enumerated. No contradictions with the annotation profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads purpose, then the flow, then Args, then Returns. Every section earns its place: the Args and Returns blocks are necessary given 0% schema coverage and no output schema, and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no parameter descriptions in the schema, the description fully specifies input constraints, the exact return shape (fields/confidence/source_text/source, account_type_hint) and the error shape. Combined with annotations covering safety, nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description carries the burden and does so well: the single parameter is base64-encoded PDF, capped at <=5 MB decoded, and constrained to non-scanned/machine-readable PDFs. An agent has everything needed to construct the argument correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Parse a base64-encoded PDF bank/credit-card/loan statement') and immediately scopes the effect ('return the extracted fields WITHOUT saving anything'). It also distinguishes itself from the apply step, so the agent can separate it from apply_statement without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames this as step 1 of a two-step flow whose output feeds a separate apply step, states the Premium-only requirement, and documents the when-not case (image-only/scanned PDFs aren't supported). The alternative (apply) and the precondition for using it are both named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reactivate_subscriptionSubscription: Reactivate canceled subscriptionAInspect
Undo a pending cancellation (only works while the subscription is still in its current paid period).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate this is not read-only, not destructive, and not idempotent. The description adds a temporal precondition not present in annotations, which is useful, but it does not disclose failure behavior, side effects, or authentication needs beyond that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that first states the action and then gives the critical condition. No wasted words, and the structure is ideal for quick agent scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that there are no parameters, no output schema, and annotations cover safety traits, the description provides the necessary operational context: what it does and when it works. It could mention failure modes or return expectations, but for a zero-parameter action this is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description adds no parameter details, but none are needed; the input schema already confirms an empty parameter set.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Undo a pending cancellation') on a clear resource, distinguishing it from the sibling tool cancel_subscription. It also adds a timeframe condition that clarifies exactly what operation it performs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear when-not constraint ('only works while the subscription is still in its current paid period'), which helps an agent decide if the tool is applicable. It does not explicitly name alternative tools, but the condition is a strong usage guideline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recategorize_similar_transactionsTransactions: Recategorize similar transactionsAInspect
Apply a category to ALL of the user's transactions that share an exact
description and type (case-insensitive description match).
Mirrors the web "recategorize similar" prompt: after fixing one
transaction's category, sweep every other transaction with the same
description + type into that category. Useful for cleaning up a
recurring merchant in one shot.
Args:
description: The exact transaction description to match
(case-insensitive; whole-string, not a substring).
transaction_type: 'deposit', 'withdrawal', or 'transfer' — only
transactions of this type are matched.
category_name: Category to assign (case-insensitive; user-owned or
system).
exclude_transaction_id: Optional transaction id to skip (e.g. the
one the user already categorized).
Returns:
``{"success": True, "updated": N, "category": ..., "category_id":
...}`` (``updated`` may be 0 if nothing else matched), or
``{"error": "..."}`` on validation failure.
| Name | Required | Description | Default |
|---|---|---|---|
| description | Yes | ||
| category_name | Yes | ||
| transaction_type | Yes | ||
| exclude_transaction_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations providing no safety signals, the description carries the behavioral burden and does so well: it warns that ALL matching transactions are affected, defines case-insensitive exact matching versus substring, and documents the return shape including updated=0 and validation errors. It could add explicit caution about overwriting existing categories, but the sweep language already implies it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is organized into purpose, usage context, Args, and Returns with no filler. Each sentence earns its place, and the most decision-relevant fact (ALL matching transactions) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter mutation with no output schema, the description is complete: matching semantics, allowed values, optional exclusion, success and error return contracts are all present. Nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the description compensates fully: description must be whole-string, transaction_type lists the three allowed values, category_name may be user-owned or system and is case-insensitive, and exclude_transaction_id is optional with a concrete example. Every parameter receives semantics beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action ('Apply a category to ALL transactions') and a precise matching rule (exact description and type, case-insensitive), which clearly distinguishes it from generic categorization tools. It also positions itself as the bulk sweep companion to a single-category fix.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the intended invocation context ('after fixing one transaction's category') and a motivating use case ('cleaning up a recurring merchant in one shot'). It does not explicitly name alternatives or when-not-to-use, but the context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reject_debt_payment_matchDebt Payments: Reject debt-payment matchAInspect
Reject a pending debt payment match. The source transaction is left
unchanged; no debt-account transaction is created.
Args:
debt_payment_id: ID of the pending DebtPayment record to reject
Returns:
Updated debt payment details
| Name | Required | Description | Default |
|---|---|---|---|
| debt_payment_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Although annotations already mark this as mutating (readOnlyHint=false) and non-destructive (destructiveHint=false), the description adds meaningful side-effect detail: the source transaction remains unchanged and no debt-account transaction is created. It does not discuss reversibility or permissions, but for this simple mutation the added transparency is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly structured: a one-sentence action, a one-sentence behavioral qualifier, and compact Args/Returns entries. No filler sentences are present, and each line adds value beyond the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema, the description covers purpose, side effects, parameter meaning, and a basic return expectation. It is almost complete; it only lacks explicit guidance about how reject relates to confirm/suggest siblings, which is already partially covered by purpose clarity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only the name and type of debt_payment_id (0% description coverage), so the description must carry the full semantic load. The Args block does exactly that: it defines debt_payment_id as the ID of the pending DebtPayment record to reject, fully disambiguating the single parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with the specific action 'Reject a pending debt payment match' and then clarifies the key behavioral scope: the source transaction is left unchanged and no debt-account transaction is created. This makes the tool's purpose clear and distinguishes it from related siblings like confirm_debt_payment_match by the explicit 'reject' semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the verb and object: an agent should call this to decline a pending debt payment match. However, the description does not explicitly name the alternative tools or state when to prefer reject over confirm_debt_payment_match, so the guidance is adequate but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remove_family_memberFamily: Remove family memberADestructiveInspect
Remove a member from your family group. Requires owner or admin role.
The group owner can't be removed.
Removes all shares involving the target member.
Args:
member_id: FamilyMember ID (not user ID) — as shown in the group's members list
Returns:
Confirmation on success, or error dict
| Name | Required | Description | Default |
|---|---|---|---|
| member_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the destructiveHint annotation, the description adds genuinely non-obvious behavior: the required owner/admin role, the constraint that the owner cannot be removed, and the destructive side effect that all shares involving the target member are removed. These are exactly the details an agent needs before invoking a destructive mutation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The critical routing constraint and permission requirement are front-loaded, and the Args/Returns blocks are compact. Every sentence earns its place, and the return-value note is useful given no output schema exists.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter destructive mutation with no output schema, the description covers purpose, authorization, an excluded case, a cascading side effect, parameter disambiguation, and the return shape. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description fully compensates by disambiguating member_id as a 'FamilyMember ID (not user ID)' and pointing to where it is found ('the group's members list'). This prevents a realistic and costly confusion between the family-member identifier and the user identifier.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Remove a member from your family group'), which is unambiguous and distinguishes it from a generic delete. It does not, however, explicitly differentiate itself from near siblings such as leave_family_group or revoke_family_invitation, leaving the agent to infer the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear preconditions for use ('Requires owner or admin role') and an exclusion ('The group owner can't be removed'), which tells the agent when the call will succeed or fail. It stops short of naming alternative tools for related operations, so there is no explicit routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rename_agent_chat_sessionAgent Chat: Rename agent chat sessionAInspect
Rename one of the user's AI assistant chat sessions.
Args:
session_id: The chat session's ID.
title: New title (1-100 chars after strip).
Returns:
{"success": True, "session": {...}} or {"error": "..."}.
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| session_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false and openWorldHint=false, so the safety profile is covered. The description adds the return envelope shape and the ownership scope ('the user's' sessions), which is useful, but says nothing about what happens on invalid IDs, permission failures, or the effects of replacing an existing title.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by a compact Args/Returns block with no filler. Slightly docstring-flavored and restates the parameter names, but nothing is padded or wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With only two required parameters, no output schema, and annotations covering the mutation/safety profile, the definition is nearly complete: it documents both params and discloses the success/error return shape. The remaining gap is behavioral (failure modes, ID discovery), which is minor for a simple rename.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% (params only carry generic titles like 'Title' and 'Session Id'), so the description must compensate. It does: session_id is identified as the chat session's ID and title is constrained to 1-100 characters after stripping, which is meaningful validation detail the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (rename) and a specific resource (one of the user's AI assistant chat sessions). It is unmistakable against the sibling chat-session tools (get/list/delete_agent_chat_session), which use different verbs, so no agent would confuse the operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use, when-not-to-use, or alternative guidance. It only implies that the session must already exist and belong to the user; there is no routing advice to related tools such as list_agent_chat_sessions (to find an ID) or delete_agent_chat_session.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reorder_accountsAccounts: Reorder accountsAInspect
Set the dashboard display order of the user's accounts.
Args:
account_ids: Account IDs in the desired top-to-bottom order. Every id
is required to be one of the user's active accounts. Any active account you
omit keeps its current order value, so pass the full list for a
deterministic ordering.
Returns:
``{"success": True, "reordered": N}``, or ``{"error": ...}`` if any id
isn't one of the user's active accounts.
| Name | Required | Description | Default |
|---|---|---|---|
| account_ids | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide idempotentHint=false, but the description adds important behavioral detail: omitting an active account retains its current order (non-deterministic unless full list passed), and failure occurs if any id is not active. This adds context beyond annotations, though it doesn't mention auth requirements or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action, then param constraints, then return format. Every sentence adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given one parameter, no output schema, and no enums, the description covers the essential behavior and return shape. It could mention whether the operation is atomic or what happens on partial success, but otherwise it's complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains that account_ids should be the full list in top-to-bottom order, that every id must be an active account, and that omission keeps current order. This is substantial semantic detail beyond the schema's bare array definition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (reorder/set display order) and resource (user's accounts) with precise scope ('dashboard display order'). It distinguishes itself from siblings like update_account or list_accounts by focusing solely on ordering.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implicitly covers when to use this tool (to set display order) but doesn't explicitly name alternatives or exclusions. For example, it doesn't clarify that this is not for updating account properties, which a sibling update_account does.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reply_to_support_ticketSupport: Reply to support ticketAInspect
Reply to one of the user's own support tickets.
If the ticket was resolved or closed, replying reopens it (the
service layer handles this), so the user can resume the conversation.
Writes to Zoninga's own support desk (no third-party service).
API docs: https://zoninga.com/ai/#support-tickets
Args:
ticket_id: The ID of the ticket to reply to.
body: The reply text (1-5000 chars).
Returns:
``{"success": True, "reply": {...}, "ticket": {"id", "status"}}``
(``ticket.status`` reflects any reopen), or ``{"error": "..."}``
if the ticket isn't found / isn't the caller's, or the body is
invalid.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | ||
| ticket_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare the generic write profile (readOnlyHint=false, idempotentHint=false, destructiveHint=false); the description adds the important non-obvious side effect that replying reopens resolved/closed tickets, clarifies it writes to Zoninga's own desk rather than a third party, and enumerates failure modes (ticket missing, not the caller's, invalid body).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in a single clear sentence, and the reopen/side-effect detail follows logically. The docstring-style Args/Returns blocks and API link add some bulk, but each section carries information rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the return shape including how ticket.status reflects a reopen and the error envelope, plus input constraints and error cases. An agent has everything needed to invoke and interpret a correct call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the burden, and it documents both parameters and adds a meaningful constraint (body 1-5000 chars) absent from the schema. The ticket_id note is thin, merely restating the name, which keeps it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (reply) and resource (support ticket) with an ownership scope ("one of the user's own"), which cleanly separates it from create_support_ticket, get_support_ticket, and list_support_tickets. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the use case clear (replying to an existing ticket) and adds a valuable condition: if the ticket was resolved or closed, replying reopens it. It does not explicitly name siblings like create_support_ticket as the alternative for new tickets, so it falls short of a full routing statement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
reset_onboardingOnboarding: Reset onboardingAInspect
Reset the authenticated user's onboarding state.
Covers requests like "can we restart the tour" or "reset my
setup checklist". Passing just one of the flags is fine — omitting
both is a no-op (returns success with nothing changed).
Args:
reset_tour: Set to True to replay the spotlight tour on next
login (clears ``UserPreferences.has_seen_tour``).
reset_checklist: Set to True to clear every ``checklist_*``
completion flag AND re-show the checklist widget.
Returns:
{"success": True, "reset_tour": bool, "reset_checklist": bool}
reflecting which pieces were reset.
| Name | Required | Description | Default |
|---|---|---|---|
| reset_tour | No | ||
| reset_checklist | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Although annotations cover the write/safety profile (readOnlyHint=false, destructiveHint=false), the description adds real behavioral context: it names the exact state being cleared (UserPreferences.has_seen_tour, all checklist_* flags), notes that the checklist widget is re-shown, and documents the no-op outcome. This is more than the annotations convey and there is no contradiction with them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose, then uses clearly labelled Args and Returns sections so an agent can skim to the relevant part. Slightly verbose with embedded markup and a multi-line Returns block, but every sentence carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-flag mutation with no output schema and no annotation detail, the description supplies everything needed: what each flag changes, the no-op case, and the exact return payload shape reflecting what was reset. Nothing required to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and neither parameter carries a description in the schema, so the description carries the full burden. It defines both flags precisely, including the side effects of each and the fact that passing only one is acceptable, fully compensating for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Reset the authenticated user's onboarding state') and immediately scopes it to the tour and setup checklist, which distinguishes it from adjacent tools like get_setup_status or dismiss_celebrations. No wasted framing before the core action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete user phrasings that should route here ('can we restart the tour', 'reset my setup checklist') and explains the flag combinations, including that omitting both is a no-op. It stops short of naming a sibling alternative or an explicit 'do not use this for X' exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
revoke_family_invitationFamily: Revoke family invitationAInspect
Revoke (cancel) a pending invitation YOU sent, before it's accepted.
The id is the one shown in the user's "sent" invitations — e.g.
to withdraw an invitation sent to the wrong address. Only group
owners/admins can revoke. (2026-07-19: closes the QA-sweep gap where
a sent invitation was unremovable from the agent surface.)
Args:
invitation_id: ID of the sent invitation to revoke.
Returns:
{"success": True, "email": ...} or an error when not found /
not pending / not yours to revoke.
| Name | Required | Description | Default |
|---|---|---|---|
| invitation_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), and the description adds genuinely new behavioral detail: the authorization requirement (owners/admins), plus the failure modes ('not found / not pending / not yours to revoke'). The stated return shape ({'success': True, 'email': ...}) is useful since no output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded and readable, but it carries internal changelog boilerplate ('(2026-07-19: closes the QA-sweep gap where a sent invitation was unremovable from the agent surface.)') that adds no operational value for an agent. That sentence does not earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly supplies the return value and error conditions, and it covers the permissions and lifecycle constraint (pending, pre-acceptance). Combined with the annotations, an agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter, and it does: invitation_id is defined as the ID from the user's 'sent' invitations, which disambiguates it from a member or invitee ID. It omits the integer type/format, but for a single required identifier the added meaning is substantial.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (revoke/cancel) plus the exact resource and scope: a pending invitation the caller sent, before acceptance. This cleanly distinguishes it from decline_family_invitation (received invites) and list_family_invitations without needing either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use context ('pending invitation YOU sent', 'before it's accepted', 'sent to the wrong address') and a prerequisite (only group owners/admins can revoke). It does not explicitly name a sibling tool to use instead, but the sent-vs-received framing implicitly routes the agent away from decline/accept.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_payment_linkAgent: Save payment linkAInspect
Save (or update) the external payment URL for one payment context.
Use when the user gives or corrects a link — e.g. "for child support
use https://childsupport.state.mn.us". Saving a label that already
exists (case-insensitive) UPDATES that link's URL.
Intended for URLs the user explicitly provided or confirmed, not
found, guessed, or constructed ones. https only; a bare domain like "chase.com" is normalized to
https://; javascript:/http:/IP-literal/credential-bearing URLs are
rejected.
Args:
label: The payment context name (max 100 chars), e.g. "child support".
url: The user-provided payment page URL (https).
Returns:
{"success": True, "created": bool, "payment_link": {...}} —
created is False when an existing label's URL was updated.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| label | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only give the generic safety profile; the description adds the crucial upsert semantics (an existing case-insensitive label UPDATES its URL, aligning with idempotentHint=false), the validation/normalization rules (https only, bare domain normalized, javascript:/http:/IP-literal/credential URLs rejected), and a field-length limit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the action, then usage, then constraints, then Args/Returns. Each sentence carries distinct information; nothing is repeated from structured fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-param mutation with no output schema, the description supplies the return shape ({success, created, payment_link}) and explains created=False when updating, leaving no gap an agent would need.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden: it defines label (payment context name, max 100 chars, example 'child support') and url (user-provided https payment page). No ambiguity remains.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (save/update) and resource (external payment URL for one payment context), and implicitly differentiates from delete_payment_link and list_payment_links by scoping to a single context's URL.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it (user gives or corrects a link) and when not to (URLs 'found, guessed, or constructed' rather than user-provided). Names the exact triggering scenario with an example.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_family_chat_messageFamily: Send family chat messageAInspect
Send a message in the family group chat.
Args:
content: Message text (max 500 characters)
Returns:
The sent message details
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are all false (non-read-only, non-idempotent, non-destructive), so the description carries some burden. It adds that the message text has a 500-character maximum and that the call returns 'the sent message details,' which is useful context. However, it does not disclose side effects like notifications to family members, error conditions, or the exact structure of the returned details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single clear sentence plus a compact Args/Returns block. Every word adds value, and the purpose is front-loaded; there is no redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter mutation with no output schema, the description gives the basic action and a return hint, but it omits prerequisites such as belonging to a family group, potential error conditions, and what fields the 'message details' actually include. This information would help an agent use the returned value for subsequent edit/delete operations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema provides only a type and title for 'content' with no description, and schema coverage is 0%. The description compensates fully by stating 'Message text (max 500 characters),' which defines both the meaning and a constraint of the single parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Send') and names a concrete resource ('a message in the family group chat'). This clearly distinguishes the tool from sibling actions like edit_family_chat_message, delete_family_chat_message, and list_family_chat_messages, and it is not a tautology of the title.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives such as edit_family_chat_message or delete_family_chat_message. It only states the action itself, leaving the selection among siblings entirely to inference from the tool name and schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
send_family_invitationFamily: Send family invitationAInspect
Send a family group invitation to an email address. Requires the group owner or admin role.
Args:
email: Email address to invite
message: Optional personal message
Returns:
Confirmation of invitation sent
| Name | Required | Description | Default |
|---|---|---|---|
| Yes | |||
| message | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, so the safety profile is covered. The description adds the role/authorization requirement and the success outcome ('Confirmation of invitation sent'), which are genuine behavioral facts beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the action in the first sentence, which is good. The Args/Returns block duplicates the schema's parameter listing somewhat, but the Returns line earns its place since there is no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation with annotations covering the safety profile and no output schema, the description supplies the action, the authorization requirement, both parameter meanings, and an outcome. Only explicit when-to-use guidance is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the params and it does: 'email: Email address to invite' and 'message: Optional personal message'. It clarifies the message is optional and personal, adding meaning beyond the bare property names, though the detail is shallow.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Send a family group invitation') with the delivery target ('to an email address'). The action is unambiguously distinct from the family invite siblings (accept/decline/revoke/list_family_invitations), so an agent can route to it without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Adds the precondition 'Requires the group owner or admin role', which tells the agent the caller's authority context. However, it never says when to prefer this over related tools (e.g., create_family_group, revoke_family_invitation) or when not to send, so usage is only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_checkoutSubscription: Start checkout (returns Stripe URL)AInspect
Create a Stripe Checkout session and return the URL for the user
to open in a browser to complete payment.
Args:
plan_id: SubscriptionPlan.id to subscribe to (a paid plan only)
Returns:
``{"checkout_url": "...", "plan_name": "..."}`` or ``{"error": ...}``.
The URL opens Stripe Checkout in the user's browser; checkout
cannot be driven programmatically. A user with a live paid
subscription gets an error pointing at the billing portal; a
user whose Premium is comped (or free in beta) still gets the URL
plus ``already_premium_note``, which explains they already have
Premium access.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations by disclosing the side effect (a Stripe session is created and a URL returned), the hard constraint that checkout cannot be driven programmatically, and three distinct outcome branches: normal URL, error for live paid subscribers, and URL plus already_premium_note for comped/beta Premium users. Annotations only state the generic non-read-only, non-destructive profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action and the return, then uses clearly labeled Args/Returns blocks for the details. No sentence is filler; even the edge-case notes earn their place by preventing wrong invocations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description compensates by spelling out the return keys (checkout_url, plan_name, error, already_premium_note) and the conditions producing each. Combined with the paid-plan constraint and the browser caveat, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter meaning — and it does: plan_id is a SubscriptionPlan.id, restricted to paid plans. It does not say where to obtain a valid plan id (e.g. list_subscription_plans), which is the one remaining gap for a 1-param tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a Stripe Checkout session and return the URL') and makes the scope explicit — the URL is for a browser, not programmatic use. This clearly separates it from siblings like open_billing_portal, list_subscription_plans, or cancel_subscription.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives substantive context: plan_id must be a 'paid plan only', and a user with a live paid subscription will get an error pointing at the billing portal — effectively steering existing subscribers elsewhere. It stops short of naming the alternative tools (open_billing_portal, list_subscription_plans) explicitly, so it is strong but not fully routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_debt_payment_patternDebt Payments: Suggest a debt-payment description patternARead-onlyIdempotentInspect
Suggest a description pattern for a new debt-payment rule.
Scans recent withdrawals on the source (asset) account for ones whose
description contains the debt account's institution name — the same
helper the web "suggest pattern" button uses. Read-only (suggests, does
not create a rule). The returned ``pattern`` is in the format
debt-payment rules accept.
Args:
source_account_id: The asset account the payment comes from.
debt_account_id: The debt account being paid.
Returns:
``{"pattern": str|None, "sample_transactions": [{id, description,
amount, date}]}`` — ``pattern`` is None when the debt account has no
institution set or no matching withdrawals were found. Returns
``{"error": "..."}`` if either account isn't the caller's.
| Name | Required | Description | Default |
|---|---|---|---|
| debt_account_id | Yes | ||
| source_account_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint=true, idempotentHint=true, destructiveHint=false), the description discloses the scanning mechanism, that no rule is created, the exact pattern format debt-payment rules accept, that pattern is None when no institution is set or no matches exist, and that an error is returned for non-owned accounts.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then structured Args/Returns sections. Every element earns its place, though the Returns block is somewhat verbose — justified here because no output schema exists to carry return-shape information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the return contract (pattern plus sample_transactions fields) and both edge cases (None pattern, error on unauthorized accounts). Nothing an agent needs to invoke and interpret the result is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full burden. The Args section fully compensates by defining source_account_id as the asset/paying account and debt_account_id as the debt account being paid, which the bare integer schema does not convey.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb+resource: 'Suggest a description pattern for a new debt-payment rule.' It distinguishes itself from siblings like create_debt_payment_rule by clarifying it 'suggests, does not create a rule', so an agent can route correctly without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the operational context (scans recent withdrawals on the source account for matching institution names, mirroring the web 'suggest pattern' button) and clarifies it is read-only. The implicit alternative is create_debt_payment_rule, but the exclusion is only hinted ('does not create a rule') rather than explicitly naming the sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
switch_accountProfile: Switch active account contextAInspect
Revoke the current session and force re-authorization.
Relevant when the session is authenticated as the wrong user. After
this runs, the connector prompts for re-authorization, and the
'Use a different account' link on the consent page signs in as a
different user.
Returns:
Dictionary with status and instructions
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly=false, destructive=false, idempotent=false, openWorld=false). The description adds real behavioral context beyond them: the session is revoked, re-authorization is triggered, and the consent-page link signs in a different user, plus a stated return shape. This is meaningful disclosure for a mutation-adjacent tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action, then the usage condition, then the outcome, and finally the return value. It is slightly padded but every sentence carries information; nothing is wasted or redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter tool with no output schema, the description supplies the trigger condition, the side effects, and an expected return ('status and instructions'). An agent has enough to invoke it correctly, though side-effect reversibility and post-switch state are not fully elaborated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is no parameter meaning to add or omit, and the schema is trivially complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (revoke the current session and force re-authorization) that is clearly distinct from sibling tools like update_user_profile or get_user_profile. It explains the mechanism behind 'switch_account' rather than restating the name. It doesn't explicitly contrast with any sibling, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear when-to-use condition: 'Relevant when the session is authenticated as the wrong user.' That is concrete context for selection, but no alternatives or when-not-to-use guidance is offered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
sync_plaid_itemPlaid: Sync linked bankAInspect
Trigger an on-demand transaction sync for one Plaid item.
Same code path as the daily cron — pulls new transactions since
the last cursor and applies them to the matched accounts.
Args:
item_id: The PlaidItem.id to sync
Returns:
Sync result counters (added/modified/removed transaction counts).
| Name | Required | Description | Default |
|---|---|---|---|
| item_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description reveals behavior beyond the sparse annotations: it 'pulls new transactions since the last cursor and applies them to the matched accounts' and uses the same code path as the daily cron. This conveys side effects and the nature of the operation without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-organized: two lead sentences plus minimal Args/Returns sections. Key information is front-loaded and every line contributes meaningfully.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter tool with no output schema, the description covers the trigger, side effects, parameter semantics, and return values. Nothing critical is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the Args section fully resolves the schema's vague 'Item Id' by stating item_id is 'The PlaidItem.id to sync.' The single parameter's meaning is completely specified in the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Trigger an on-demand transaction sync') on a specific resource ('one Plaid item'), clearly distinguishing it from read/list/delete siblings like get_plaid_item, list_plaid_items, and unlink_plaid_item. The title reinforces the same operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly frames the tool as the on-demand counterpart to the daily cron, giving clear context for when it should be used. It does not name alternatives or exclusion conditions, but the single-purpose nature and cron comparison make the intended usage clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
undo_transactionTransactions: Undo transactionAInspect
Restore a soft-deleted transaction created manually. Only works
for transactions deleted (via MCP or the web UI) within the last
7 days. Plaid-managed transactions cannot be
undone (they are hard-deleted on their original delete).
Args:
transaction_id: The ID of the soft-deleted transaction to restore
Returns:
The restored transaction details, or an error dict
| Name | Required | Description | Default |
|---|---|---|---|
| transaction_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds substantial traits beyond the annotations: a 7-day recovery window, an explicit carve-out for Plaid-managed transactions, and the fact that it can fail with an error dict. It is consistent with the annotations (readOnlyHint=false for a mutation) and enriches them rather than contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and constraints are front-loaded, and the Args/Returns block is compact. The structured style is slightly mechanical but every line carries information and nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter mutation with no output schema, the description covers what it does, the preconditions, the failure mode, and the return value ('restored transaction details, or an error dict'). Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the parameter. It does: 'The ID of the soft-deleted transaction to restore' clarifies that the input must reference a soft-deleted record, which meaningfully constrains valid values beyond the bare integer type.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource ('Restore a soft-deleted transaction') and immediately scopes it against the broader transaction family (manually created, soft-deleted, not Plaid-managed). An agent can distinguish it from delete_transaction/update_transaction without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives strong preconditions: only soft-deleted transactions, only ones deleted within the last 7 days, and explicitly not Plaid-managed transactions ('hard-deleted on their original delete'). This is clear when-to-use/when-not context, though it doesn't name the inverse sibling (delete_transaction) as an alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
unlink_plaid_itemPlaid: Unlink bankADestructiveInspect
Unlink a Plaid item (soft-deactivate) and revoke the upstream
access token at Plaid.
Linked accounts are NOT deleted — they're detached from Plaid and
become manual accounts (the user keeps their transaction history).
Args:
item_id: The PlaidItem.id to unlink
Returns:
Status with the institution name and detached-account count.
| Name | Required | Description | Default |
|---|---|---|---|
| item_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses critical behavioral traits: the operation is a soft-deactivate, it revokes the upstream access token, linked accounts are not deleted, they become manual accounts, and transaction history is preserved. This directly enriches the destructiveHint=true annotation without contradicting it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured, with the core action in the first sentence, important side-effect clarification immediately after, and clearly labeled Args and Returns sections. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, destructive tool with no output schema, the description is complete: it explains the input, the side effects, the non-destructive outcome for accounts, and the return status content. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does by explaining that item_id is 'The PlaidItem.id to unlink,' giving the integer parameter a clear semantic identity and linking it to the entity being unlinked.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Unlink a Plaid item (soft-deactivate) and revoke the upstream access token at Plaid.' It also clarifies the key distinction from deletion by explaining that linked accounts become manual accounts, which distinguishes it from sibling tools like delete_account or sync_plaid_item.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case: when a user wants to unlink a Plaid bank connection while preserving account data. However, it does not explicitly state when to prefer this over related tools like sync_plaid_item or delete_account, nor does it mention exclusions, so the guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_accountAccounts: Update accountAInspect
Update an existing financial account.
Args:
account_id: The account ID to update
name: New account name (optional)
account_type: New account type (optional). One of: checking, savings,
cash, credit_card, loan, mortgage, student_loan, heloc, arm,
investment, other_asset, other_debt
balance: New balance amount (optional, use with caution)
institution: New institution name (optional)
currency: New currency code. App is locked to USD pre-launch — any
other value is rejected. (See docs/CURRENCY_RESTORE_TODO.md.)
is_active: Set account active/inactive (optional)
interest_rate: New interest rate for debt accounts (optional)
credit_limit: New credit limit for credit cards (optional)
minimum_payment: New minimum payment for debt accounts (optional)
term_months: Original loan term in months (optional)
loan_start_date: Loan origination date YYYY-MM-DD (optional)
loan_end_date: Loan payoff target date YYYY-MM-DD (optional)
original_balance: Original loan amount when first taken out (optional)
monthly_escrow_tax: Monthly property tax escrow, mortgage only (optional)
monthly_escrow_insurance: Monthly insurance escrow, mortgage only (optional)
draw_period_months: HELOC draw-phase duration in months (optional)
draw_amount_per_month: HELOC monthly draw amount (optional)
rate_2: 2nd-term rate as percentage, HELOC/ARM (optional)
term_2_months: 2nd-term duration in months, HELOC/ARM (optional)
arm_initial_period_months: ARM fixed-rate period in months (optional)
include_in_debt_paydown: Include this debt in the payoff calculator
and dashboard payoff card (optional). Set False for a debt you
pay in full each month so it doesn't distort the plan; does NOT
affect net worth, total debt, or the balance sheet.
low_balance_alert_threshold: Per-account low-balance alert
threshold in dollars — "alert me when this account drops below
$X" (checking/savings/cash). Overrides the monitoring plan's
global threshold for this account only.
clear_low_balance_threshold: Set True to remove a per-account
threshold and fall back to the plan's global default.
mute_low_balance_alerts: True = no low-balance alerts on this
account (daily brief + notifications) — "stop reminding me,
I know this account is low". False = re-enable alerts.
institution_login_url: Where the user logs in at this institution
(https only, e.g. "https://chase.com"; bare domains are
normalized to https). An empty string removes the link. Intended
for a URL the user provided, not a guessed or invented one.
Returns:
Updated account details
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| rate_2 | No | ||
| balance | No | ||
| currency | No | ||
| is_active | No | ||
| account_id | Yes | ||
| institution | No | ||
| term_months | No | ||
| account_type | No | ||
| credit_limit | No | ||
| interest_rate | No | ||
| loan_end_date | No | ||
| term_2_months | No | ||
| loan_start_date | No | ||
| minimum_payment | No | ||
| original_balance | No | ||
| draw_period_months | No | ||
| monthly_escrow_tax | No | ||
| draw_amount_per_month | No | ||
| institution_login_url | No | ||
| include_in_debt_paydown | No | ||
| mute_low_balance_alerts | No | ||
| monthly_escrow_insurance | No | ||
| arm_initial_period_months | No | ||
| clear_low_balance_threshold | No | ||
| low_balance_alert_threshold | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotations declaring a non-idempotent, non-destructive write, the description adds real behavioral context beyond them: non-USD currency is rejected pre-launch, institution_login_url is https-only and normalizes bare domains, empty string removes the link, and include_in_debt_paydown explicitly does NOT affect net worth/total debt/balance sheet. The 'use with caution' on balance is vague, and idempotency behavior isn't addressed, keeping it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, and the parameter block is dense but every line adds needed meaning given 0% schema coverage. The Python 'Args:'/'Returns:' docstring framing and an internal file reference (docs/CURRENCY_RESTORE_TODO.md) are minor formatting leaks rather than wasted content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 26-parameter mutation with no output schema and minimal annotation detail, the description covers every input precisely and states the return ('Updated account details'). The only thinness is the terse return statement and lack of any error/permission behavior, but the input surface is complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the entire burden — and it documents all 26 parameters with meaning beyond names: enum values for account_type, YYYY-MM-DD formats for loan dates, 'mortgage only' scoping for escrow fields, HELOC/ARM scoping for second-term fields, and behavioral semantics for the threshold/alert parameters (override vs global default, clear vs set). This is exactly the compensation the low coverage requires.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource: 'Update an existing financial account.' An agent immediately knows this mutates an account record. However, it does not distinguish itself from close siblings like override_account_balance, create_account, or update_account's relationship to them, so an agent must infer routing from names alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no tool-level when-to-use/when-not guidance and no named alternative, despite several plausible siblings (override_account_balance for balance edits, delete_account, reorder_accounts). The caution note on 'balance' hints at constraints but is not routing guidance. Usage must be inferred entirely from the parameter list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_budgetBudgets: Update budgetAInspect
Update an existing budget.
Args:
budget_id: The ID of the budget to update
amount: New budget amount (optional)
period: New period - "monthly" or "yearly" (optional)
rollover: Enable/disable rollover (optional, monthly budgets only)
name: New display name (optional, Phase 92). Pass an empty string to
clear the custom name and fall back to the category name.
Returns:
Updated budget details
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| amount | No | ||
| period | No | ||
| rollover | No | ||
| budget_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the neutral annotations, the description discloses meaningful behavior: it returns 'Updated budget details,' rollover is only applicable to monthly budgets, and passing an empty name clears the custom name and falls back to the category name. This gives an agent useful insight into edge-case behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the action and uses a clear Args/Returns structure. Every line adds useful parameter information; only the 'Phase 92' note is slightly tangential, but it does not prevent the definition from being compact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple partial-update tool with no output schema, the description covers the required budget_id, all optional fields, valid values, and the return value. It could note whether at least one optional field is required and clarify how period changes interact with rollover, but nothing essential is missing for a straightforward call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description documents all five parameters with purpose, optionality, valid values for period, and behavioral notes for rollover and name. This fully compensates for the empty schema and gives the agent semantics the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Update an existing budget,' which names a specific verb and resource, and the word 'existing' distinguishes it from create_budget and delete_budget. It is clear but does not explicitly call out sibling tools or state when not to use it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance is provided beyond the verb itself. The description does not direct an agent toward create_budget for new budgets, list_budgets for finding budget IDs, or delete_budget for removals, and it does not state prerequisites for calling the tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_categoryCategories: Update categoryAInspect
Update a user-created category. System categories cannot be modified.
Args:
category_id: The category ID to update
name: New category name (max 50 chars)
Returns:
Updated category details
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| category_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations (readOnlyHint=false) confirm it's a mutating operation and do not contradict the description. The description adds the key behavioral fact that system categories are protected from modification. It does not disclose other traits such as error handling, idempotency, or side effects, but given the simple nature of the tool, the added constraint provides meaningful transparency beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: two sentences plus an Args/Returns breakdown. It front-loads the core purpose and constraints, with no superfluous content. Every sentence earns its place, and the structured format aids quick parsing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple 2-parameter tool with no nested objects and no output schema, the description covers the essential context: what it does, the system category restriction, parameter meanings, and a general return type. It does not specify error cases or permission requirements, but these are not critical for a simple update tool and the existing info is sufficient for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain parameters. It does so effectively: category_id is described as 'the category ID to update' and name as 'new category name (max 50 chars)'. This adds value beyond the bare schema by clarifying purpose and validation. It omits that name is optional (though schema shows default null), but this is a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the action: 'Update a user-created category.' The verb 'update' plus the resource 'category' is specific. It also distinguishes from create/delete by saying this is for existing user-created categories, and adds the constraint that system categories cannot be modified, which sets scope apart from potentially similar tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implicitly indicates usage: update user-created categories, and explicitly warns that system categories cannot be modified. However, it does not mention alternatives (e.g., create_category for new categories, delete_category for removal) or provide explicit when-to-use vs. when-not-to-use guidance beyond the system category restriction. This leaves some inference to the agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_debt_payment_ruleDebt Payments: Update debt-payment ruleAInspect
Update the pattern or account pairing on an existing debt payment rule.
Args:
rule_id: ID of the rule to update
description_pattern: New substring to match (optional, <=200 chars)
source_account_id: New asset account (optional)
debt_account_id: New debt account (optional)
Returns:
Updated rule details
| Name | Required | Description | Default |
|---|---|---|---|
| rule_id | Yes | ||
| debt_account_id | No | ||
| source_account_id | No | ||
| description_pattern | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations show readOnlyHint=false, so the description carries the burden of explaining mutation behavior. It states that fields can be updated and that updated rule details are returned, but it does not disclose what happens when optional fields are omitted or set to null, whether a rule must exist first, or how partial updates behave. This is a meaningful gap for an update operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is tightly structured with a clear one-sentence summary followed by an Args list and Returns line. There is no filler, and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is sufficient for a basic call, but it lacks context about null/omitted parameter semantics, whether at least one optional field is required for a meaningful update, and what the updated rule details contain. With no output schema, a bit more detail about the return value would improve completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description compensates well by explaining all four parameters in plain language: rule_id identifies the rule, description_pattern is the new substring with a length constraint, and the two account IDs are labeled as new asset and debt accounts. It adds meaning and constraints not present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Update the pattern or account pairing on an existing debt payment rule.' It clearly identifies what is modified and can be distinguished from sibling tools like create_debt_payment_rule, delete_debt_payment_rule, and list_debt_payment_rules.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: it is for modifying an existing debt-payment rule, as opposed to creating or deleting one. However, it does not explicitly state when to prefer this over alternatives or mention any preconditions or exclusions, leaving the usage context mostly implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_forum_postForum: Update forum postAInspect
Update an existing forum post (author only).
Args:
post_id: The post to update.
title: New title (3-200 chars).
content: New raw markdown content (10-10000 chars).
Returns:
``{"success": True, "post": {...}}`` on success, or
``{"error": "..."}`` (not found, not the author, or validation).
| Name | Required | Description | Default |
|---|---|---|---|
| title | Yes | ||
| content | Yes | ||
| post_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=false, destructiveHint=false, and idempotentHint=false, so the description is not required to repeat those. It adds valuable behavioral context beyond the annotations: the operation fails with specific error categories (not found, not the author, or validation) and returns the updated post. This is meaningful and non-redundant.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core purpose. Each parameter is given a one-line treatment, and the return behavior is specified in two short lines. Every sentence carries information; there is no filler or repetition of schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter mutation with no output schema and no enums, the description is nearly complete: it covers purpose, constraints, format requirements, error behavior, and return shape. It could additionally note that all three parameters are required and that the operation overwrites existing title/content, but these are minor gaps given the schema's required list and the description's already strong coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does. It adds critical validation constraints absent from the schema: title must be 3-200 chars and content must be 10-10000 chars. It also clarifies that post_id is the target to update and content is raw markdown. This goes well beyond the bare schema property names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Update'), a resource ('forum post'), and a key scope constraint ('author only'). It is clearly distinguishable from siblings like create_forum_post and delete_forum_post, and the leading phrase leaves no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use the tool: to modify an existing forum post the user authored. It does not explicitly name sibling alternatives or state when not to use it (e.g., 'use create_forum_post to create a new post'), but the 'author only' constraint provides strong practical guidance and excludes common misuse cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_forum_replyForum: Update forum replyAInspect
Update an existing forum reply (author only).
Args:
reply_id: The reply to update.
content: New raw markdown content (5-10000 chars).
Returns:
``{"success": True, "reply": {...}}`` on success, or
``{"error": "..."}`` (not found, not the author, or validation).
| Name | Required | Description | Default |
|---|---|---|---|
| content | Yes | ||
| reply_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnlyHint=false, so the mutation is expected; the description adds concrete behavioral details: author-only enforcement, content length bounds (5-10000 chars), and possible error conditions (not found, not author, validation). This goes beyond what annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The main purpose is front-loaded in a single sentence, followed by compact Args and Returns sections. The format is slightly verbose with docstring-style labels, but every element adds useful information and there is no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 required parameters, no output schema, no nested objects), the description is fully self-contained: it states the permission model, parameter constraints, success return shape, and failure modes. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must carry the full meaning of both parameters. It does: reply_id is identified as 'the reply to update,' and content is specified as 'new raw markdown content (5-10000 chars),' providing both format and validation constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description starts with a specific verb and resource: 'Update an existing forum reply (author only).' It clearly distinguishes from sibling create_forum_reply, delete_forum_reply, and update_forum_post, removing ambiguity about what action applies to which object.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The '(author only)' parenthetical provides a clear eligibility constraint, telling the agent this tool cannot be used to modify others' replies. While it doesn't explicitly name alternatives, the sibling context and the 'existing' wording imply this is for editing an already-created reply, not creating or deleting.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_goalGoals: Update goalAInspect
Update an existing financial goal.
Args:
goal_id: The ID of the goal to update
name: New goal name (optional)
target_amount: New target amount (optional, may auto-complete if <= current amount)
target_date: New target date in YYYY-MM-DD format (optional)
status: New status - active, completed, paused, cancelled (optional)
color: New hex color code (optional)
notes: New notes (optional)
priority: New priority level (optional)
Returns:
Updated goal details
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| color | No | ||
| notes | No | ||
| status | No | ||
| goal_id | Yes | ||
| priority | No | ||
| target_date | No | ||
| target_amount | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations provide no read-only, idempotency, or destructiveness hints, so the description carries the burden. The description discloses useful behavior: target_amount 'may auto-complete if <= current amount' and status has specific allowed values, plus a Returns line. However, it leaves important mutation semantics unstated, such as whether updates are partial, whether null clears fields, or whether status transitions are validated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a compact Google-style docstring with a one-line purpose, a dense parameter list, and a Returns line. Every section earns its place, especially because the schema itself contains no parameter descriptions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter mutation tool with no annotations and no output schema, the description covers all parameters, the return value, and a notable side effect. It is still missing explicit partial-update semantics and error/precondition behavior, which would make it fully complete for an agent invoking it correctly without extra inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description compensates by documenting all 8 parameters. It adds format details for target_date (YYYY-MM-DD), enumerates status choices, specifies hex color format, and highlights target_amount auto-completion. It is slightly weaker on priority type/range and null semantics, but still provides substantial meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a precise verb and resource: 'Update an existing financial goal.' It is clearly distinguishable from sibling tools like create_goal, delete_goal, and list_goals, which operate on the same resource with different actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'existing financial goal' implies this tool is for modifying goals that already exist, but the description never explicitly says when to prefer it over create_goal, delete_goal, or contribute_to_goal. There are no exclusions or alternative-name mentions, only an implied usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_hard_assetEstate: Update hard assetCInspect
Update a hard asset's details.
Args:
asset_id: ID of the hard asset to update
name: New name (optional)
current_value: Updated current value (optional)
description: Updated description (optional)
location: Updated location (optional)
Returns:
Updated hard asset details.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| asset_id | Yes | ||
| location | No | ||
| description | No | ||
| current_value | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations only provide hints (readOnlyHint=false, destructiveHint=false), and the description does not add meaningful behavioral context beyond stating it updates and returns details. It does not disclose partial-update semantics, failure modes, permissions, or side effects, which would be valuable for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured with a lead sentence, an Args section, and a Returns section. Each line serves a purpose and there is no redundant or off-topic content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 5-parameter mutation tool with no output schema and no schema descriptions, the description is too thin. It omits details about partial updates, error handling, permissions, and the exact structure of the returned updated asset, leaving gaps an agent must resolve elsewhere.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description must compensate, and it does provide short definitions for each parameter like 'New name' and 'Updated current value'. However, these largely paraphrase the parameter names and do not explain formats, constraints, units, or currency expectations, offering only minimal added meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Update') and the resource ('a hard asset's details'), so an agent can identify the tool's purpose. However, it does not explicitly differentiate itself from sibling tools like create_hard_asset or delete_hard_asset, relying on the name rather than explicit contrast.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, such as create_hard_asset for new assets or list_hard_assets for viewing. No prerequisites, such as the asset existing or ownership requirements, are mentioned, leaving the agent to infer usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_milestoneMilestones: Update milestoneBInspect
Update a milestone's name, target amount, or celebration message.
Args:
milestone_id: The milestone ID to update
name: New name (optional)
target_amount: New target (> 0 and <= parent goal's target)
celebration_message: New celebration message (optional)
Returns:
Updated milestone details
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| milestone_id | Yes | ||
| target_amount | No | ||
| celebration_message | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, openWorldHint=false, so the safety profile is already covered. The description adds a useful constraint (target_amount must be > 0 and <= parent goal's target), which is behavioral context beyond annotations. However, it doesn't say whether partial updates replace or merge values, or whether missing fields are left unchanged.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Well-structured with explicit Args and Returns sections, front-loaded with the core action. Slightly verbose for a simple update but every line is relevant. The triple-quoted formatting is a bit rough but content is not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-param mutation tool with no output schema and 0% schema description coverage, the description covers the basics but leaves gaps: no return value detail beyond 'Updated milestone details', no merge/replace semantics, no error conditions for target_amount exceeding parent. Annotations carry safety, but behavioral completeness is only adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It does list each parameter and provides a constraint for target_amount, which adds meaning. But name and celebration_message get only 'optional' with no constraints, and milestone_id gets no format hints. Baseline is 3 given the partial compensation for low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (milestone), and enumerates the mutable fields. It does not differentiate from sibling tools like create_milestone or delete_milestone, but the verb+resource alone makes the operation type clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, no mention of alternatives. The only hint is that fields are optional, which tells an agent it can do partial updates, but there is no explicit routing to update_goal or create_milestone for related needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_monitoring_planAgent: Update monitoring planAInspect
Update (and optionally activate) the user's monitoring plan.
Only the fields passed are changed. Part of onboarding: activate=True
turns hands-off monitoring on and is intended to follow the user's
explicit confirmation in the conversation.
The plan drives notifications and Tier-2 proposals only — it moves
no money.
Args:
check_in_frequency: daily / weekly / monthly / off.
watch_*: toggle individual watchers.
low_balance_threshold: alert when an account drops below this ($).
surplus_buffer: keep at least this much in checking ($).
surplus_minimum_sweep: minimum sweep size proposed ($); smaller
sweeps aren't proposed.
watch_cashflow_shortfall / watch_anomalies: forecast + anomaly watchers.
watch_new_recurring / watch_duplicates / watch_quiet_account:
elder-watch trio — new recurring charge inference, possible
double-posted charges, and spending on a normally-quiet
account (all informational; nothing moves money).
activate: set True (with the user's explicit OK) to approve the plan
and start the sentinel.
Returns:
{"success": True, "plan": {...}} or {"error": ...}.
| Name | Required | Description | Default |
|---|---|---|---|
| activate | No | ||
| watch_surplus | No | ||
| surplus_buffer | No | ||
| watch_anomalies | No | ||
| watch_duplicates | No | ||
| watch_low_balance | No | ||
| check_in_frequency | No | ||
| watch_bill_coverage | No | ||
| watch_new_recurring | No | ||
| watch_quiet_account | No | ||
| low_balance_threshold | No | ||
| surplus_minimum_sweep | No | ||
| watch_cashflow_shortfall | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, idempotentHint=false. The description adds meaningful context beyond these: only passed fields change (patch semantics), the plan 'moves no money', and watchers are informational. This reassures the agent about side effects, though auth/permission requirements are not discussed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose before the Args/Returns detail, and every section adds value. It is longer than average but justified by 13 undocumented parameters; the docstring structure is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter mutation tool with no output schema, the description covers purpose, partial-update behavior, activation flow, per-field meaning, and even the return shape ({success, plan} / {error}). An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 13 parameters and 0% schema coverage, the description fully carries the burden: it explains check_in_frequency's allowed values (daily/weekly/monthly/off), the threshold semantics for low_balance_threshold, surplus_buffer and surplus_minimum_sweep, and annotates each watcher group. This adds substantial meaning absent from the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Update (and optionally activate) the user's monitoring plan') and clarifies the domain scope by naming what it affects ('drives notifications and Tier-2 proposals only'). An agent can distinguish this from get_monitoring_plan without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear context: 'Part of onboarding' and activate=True 'is intended to follow the user's explicit confirmation.' This tells the agent when the activate flag is appropriate. No explicit alternative tool is named, but the usage context is concrete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_recurring_transactionRecurring Transactions: Update recurring templateAInspect
Update a recurring transaction template.
Only the fields you pass are updated. Past spawned instances are
left alone — the change applies to future spawns only.
Args:
template_id: The template Transaction.id
amount: New amount (optional)
description: New description (optional)
frequency: New frequency (optional)
end_date: New end date as YYYY-MM-DD, or empty string to clear (optional)
category_id: New category id, or 0/null to clear (optional)
| Name | Required | Description | Default |
|---|---|---|---|
| amount | No | ||
| end_date | No | ||
| frequency | No | ||
| category_id | No | ||
| description | No | ||
| template_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations, the description discloses important behavior: only supplied fields are updated, past spawned instances are left alone, and changes apply only to future spawns. This is genuinely valuable behavioral context for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core behavior is stated in the first two sentences, with no filler, and the parameter list is compact and directly useful for a six-parameter tool. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a partial-update tool with no output schema, the description is self-contained: it covers all parameters, the scope of mutation, and special clearing rules. No essential information for invoking the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the description fully compensates: it explains every parameter's purpose, optionality, and special clear semantics, including YYYY-MM-DD for end_date, empty string to clear, and 0/null to clear category_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear, specific action and resource: 'Update a recurring transaction template.' This distinguishes it from update_transaction and create/delete_recurring_transaction by emphasizing the template and its future-spawn scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context by explaining partial updates and that past spawned instances are unaffected, so an agent knows when this operation fits. It does not explicitly name alternatives or exclusion conditions, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_scenarioScenarios: Update a what-if scenarioAInspect
Update a what-if scenario (only the fields you pass change).
When ``items`` is omitted the existing changes are kept; pass a full
new list to replace them. A committed scenario can't be edited.
Item dict shapes match scenario creation (add_debt / add_asset / remove).
Args:
scenario_id: The scenario to update.
name / description / strategy / extra_monthly: Optional new values.
items: Optional full replacement list of change dicts.
Returns:
{scenario: {...}, analysis: {...}} or {"error": "..."}.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| items | No | ||
| strategy | No | ||
| description | No | ||
| scenario_id | Yes | ||
| extra_monthly | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (writable, non-idempotent, non-destructive). The description adds real context beyond that: partial-update semantics, replacement-vs-keep behavior for items, and the committed-scenario immutability constraint. Return shape is also disclosed in the absence of an output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the key mutation semantics and immutability rule, then structured Args/Returns sections. Every sentence earns its place, with only mild redundancy between the prose and the Args block.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param mutation tool with 0% schema coverage and no output schema, the description is complete enough: it covers partial-update behavior, the committed constraint, item shape, and the return value. Nothing prevents correct invocation, though deeper detail on the item dict schema would help.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry the load, and it does list every parameter with meaning: name/description/strategy/extra_monthly as optional new values, items as a full replacement list whose dict shapes mirror scenario creation (add_debt/add_asset/remove). The complex items parameter is adequately explained; only granular per-field detail for the scalar args is thin.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (what-if scenario), and the parenthetical 'only the fields you pass change' immediately clarifies the mutation semantics. It does not explicitly name siblings like create_scenario or commit_scenario, but the 'what-if scenario' scope is specific enough to distinguish it from the other scenario tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear when-not rule ('A committed scenario can't be edited') and explains the branch between omitting items (keep existing) and passing a full list (replace). These are the two decision points an agent faces, though it doesn't explicitly route to alternatives like commit_scenario or get_scenario.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_tagTags: Update tagAInspect
Update a tag's name and/or color. Renaming preserves all transaction associations.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | ||
| color | No | ||
| tag_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds a valuable behavioral guarantee beyond annotations: 'Renaming preserves all transaction associations.' This clarifies that a rename is non-destructive to related data, which annotations (readOnlyHint, destructiveHint) do not fully convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler. The core action is front-loaded, and the important behavioral note about preserving associations is included without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple update tool with no output schema axis, this description adequately covers what the tool does, what can be changed, and a key side effect. It omits details like return values, but those are not expected given the simplicity of the operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must compensate. It names two fields (name, color) and clarifies that either or both can be updated ('and/or'). However, tag_id (the required parameter) is not explained, though its role is inferable from the name.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Update a tag's name and/or color') with a specific resource and fields. It does not explicitly name alternatives like create_tag or list_tags, but the verb and resource make the tool's scope unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied: this tool is for modifying an existing tag's name or color. No explicit when-to-use or when-not-to-use guidance is provided, nor are alternatives mentioned, but the description is clear enough for an agent to infer the appropriate scenario.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_transactionTransactions: Update transactionAInspect
Update an existing transaction.
Args:
transaction_id: The ID of the transaction to update
description: New description (optional)
amount: New amount as positive number (optional, will adjust account balance)
category_name: New category name, case-insensitive (optional)
transaction_date: New date in YYYY-MM-DD format (optional)
tag_names: List of tag names to assign (optional, replaces existing tags)
Returns:
Updated transaction details with new account balance
| Name | Required | Description | Default |
|---|---|---|---|
| amount | No | ||
| tag_names | No | ||
| description | No | ||
| category_name | No | ||
| transaction_id | Yes | ||
| transaction_date | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only provide basic hints (readOnlyHint=false, destructiveHint=false). The description adds meaningful behavioral context by stating that setting amount 'will adjust account balance' and that tag_names 'replaces existing tags'. These side effects go beyond what the annotations convey and are critical for an agent to anticipate the consequences of a call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: a one-line purpose, a bulleted Args list, and a Returns line. Every sentence earns its place, with the most important information (the action) front-loaded and parameter details concisely formatted. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description clearly explains the return value ('Updated transaction details with new account balance'). It covers all six parameters, their constraints, and side effects. While it doesn't mention error conditions or prerequisites like 'transaction must exist', these are common for update tools and the given context is sufficient for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry all parameter meaning. It does so thoroughly: amount is 'positive number', category_name is 'case-insensitive', transaction_date has an explicit 'YYYY-MM-DD' format, and tag_names explicitly replaces existing tags. This is exactly the type of semantic enrichment the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Update an existing transaction', a specific verb and resource that clearly distinguishes it from other update_* tools for accounts, budgets, categories, etc. The list of updatable fields reinforces the scope. No ambiguity about which entity is affected.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance about when to use this tool versus alternatives like recategorize_similar_transactions, bulk_categorize_transactions, or undo_transaction. The description simply states what it does without providing context on the appropriate conditions or excluding cases. This leaves the agent to infer usage from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_user_preferencesProfile: Update user preferencesAInspect
Update user preferences. Only provided fields are changed.
Args:
theme: Color theme (light, dark, emerald_banking, frost_glass, carbon)
currency: Currency code. USD only — multi-currency is disabled
pre-launch, so any other value is rejected (matches the web
form + REST serializer, which omit/lock the field).
notify_budget_alerts: Enable budget alert notifications
notify_bill_reminders: Enable bill reminder notifications
notify_large_transactions: Enable large transaction notifications
notify_weekly_summary: Enable weekly summary emails
notify_balance_alerts: Enable Plaid balance-mismatch alert
notifications (off by default — the dashboard badge is the
default signal)
notify_debt_payments: Enable the monthly debt-strategy payment
reminder (requires a saved debt_strategy to fire)
notify_daily_brief_email: Email the agent's morning brief to the
user's inbox (the brief appears in the app thread either
way; this toggles the email copy only)
notify_family_alerts: Receive in-app alerts + the family brief
about accounts a family member shared with the user
(elder-care; on by default, receiver-side opt-out)
notify_family_brief_email: Email the family brief (off by
default; alerts keep appearing in the app either way)
show_profile_picture: Show profile picture in forum
show_real_name: Show real name in forum
show_level: Show level badge in forum
show_badges: Show earned badges in forum
debt_strategy: Save the user's debt-paydown "my plan" strategy — one
of minimum_payments, snowball, variable_snowball, highest_rate,
highest_balance, highest_payment, cashflow_index, npv,
max_interest_savings. Pass "" to clear it. This drives the
dashboard debt-payoff card + the strategy reminder. Free plans are
limited to the lower-tier strategies (Premium-only ones rejected).
debt_extra_monthly_payment: Extra $/month applied on top of minimums
in the paydown plan (0 or positive).
Returns:
Updated preferences (including debt_strategy + debt_extra_monthly_payment)
| Name | Required | Description | Default |
|---|---|---|---|
| theme | No | ||
| currency | No | ||
| show_level | No | ||
| show_badges | No | ||
| debt_strategy | No | ||
| show_real_name | No | ||
| notify_budget_alerts | No | ||
| notify_debt_payments | No | ||
| notify_family_alerts | No | ||
| show_profile_picture | No | ||
| notify_balance_alerts | No | ||
| notify_bill_reminders | No | ||
| notify_weekly_summary | No | ||
| notify_daily_brief_email | No | ||
| notify_family_brief_email | No | ||
| notify_large_transactions | No | ||
| debt_extra_monthly_payment | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false and destructiveHint=false, so the description's job is to add context — and it does: partial updates, currency is locked to USD pre-launch, debt_strategy accepts "" to clear, free plans reject Premium strategies, and several notify_* flags carry defaults (e.g. notify_balance_alerts off by default). Auth requirements and rate limits remain unstated.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The lead sentence is front-loaded and the Args block is scannable; almost every line adds distinct information. Some parenthetical asides (e.g. the dashboard-badge note on balance alerts) run longer than needed, but nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-parameter mutation tool with no output schema, the description covers every field's meaning, the constraints that cause rejection, defaults, and what the call returns. An agent can invoke this correctly without consulting anything else.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the full parameter burden for 17 fields, and it delivers: enumerations (theme colors, debt_strategy values), accepted value ranges, clearing semantics, and plan-gating per field. This is exactly the compensation the 0% coverage demands.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (update user preferences) and adds the key partial-update contract: 'Only provided fields are changed.' That distinguishes it from sibling getters like get_user_preferences and from the profile-scoped update_user_profile, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the partial-update semantics and the per-field notes, but there is no explicit when-to-use this vs update_user_profile, and no stated prerequisites or auth context for a mutation tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_user_profileProfile: Update user profileAInspect
Update the authenticated user's profile fields.
Only the fields you pass are changed. Email is NOT editable here
(changing email requires re-verification on the web). Profile
picture upload is not exposed over MCP — it's web-only.
Args:
first_name: New first name (max 150 chars)
last_name: New last name (max 150 chars)
bio: Short forum bio (max 500 chars)
forum_display_name: Forum display name (max 50 chars). Blank
falls back to a masked email in the forum.
state: Two-letter US state code (e.g. 'CA'). '' clears the field.
Returns:
Updated profile details
| Name | Required | Description | Default |
|---|---|---|---|
| bio | No | ||
| state | No | ||
| last_name | No | ||
| first_name | No | ||
| forum_display_name | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
It discloses partial-update semantics ('Only the fields you pass are changed'), email and picture exclusions, the blank fallback for forum_display_name, and that an empty state clears the field. This adds behavioral detail beyond the annotations, though it does not describe exact response or failure behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: a one-line summary, key exclusions, a bulleted Args list, and a Returns line. Every sentence adds signal and the most important usage constraints come first.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter mutation with no output schema and no per-parameter schema descriptions, the description provides all essential call guidance: field semantics, constraints, clearing behavior, and exclusions. No critical information is missing for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description carries the full burden for parameter understanding. It documents each of the five parameters with max lengths, formats, and special behaviors (e.g., '' clears state, blank forum_display_name falls back to masked email).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Update the authenticated user's profile fields,' a specific verb and resource, then enumerates the editable fields. It also marks email and profile-picture upload as out of scope, differentiating it from web-only profile actions and nearby update_* siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states that email is not editable here and profile-picture upload is not exposed over MCP, giving clear boundary guidance. It does not explicitly name alternative tools such as update_user_preferences or get_user_profile, so some when-to-use routing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
179 tool updates
- First observed
accept_family_invitation - First observed
accept_plaid_balance - First observed
activate_theme - First observed
apply_statement - First observed
bulk_categorize_transactions - First observed
bulk_delete_transactions - First observed
cancel_subscription - First observed
check_payment_coverage - First observed
commit_scenario - First observed
confirm_debt_payment_match - First observed
contribute_to_goal - First observed
create_account - First observed
create_budget - First observed
create_category - First observed
create_debt_payment_rule - First observed
create_family_group - First observed
create_forum_post - First observed
create_forum_reply - First observed
create_goal - First observed
create_goal_plan - First observed
create_hard_asset - First observed
create_milestone - First observed
create_recurring_transaction - First observed
create_reminder - First observed
create_scenario - First observed
create_spending_challenge - First observed
create_support_ticket - First observed
create_tag - First observed
create_transaction - First observed
decline_family_invitation - First observed
delete_account - First observed
delete_agent_chat_session - First observed
delete_budget - First observed
delete_category - First observed
delete_debt_payment_rule - First observed
delete_family_chat_message - First observed
delete_forum_post - First observed
delete_forum_reply - First observed
delete_goal - First observed
delete_hard_asset - First observed
delete_milestone - First observed
delete_notification - First observed
delete_payment_link - First observed
delete_recurring_transaction - First observed
delete_reminder - First observed
delete_scenario - First observed
delete_tag - First observed
delete_transaction - First observed
dismiss_budget_recommendation - First observed
dismiss_celebrations - First observed
dismiss_payment_warning - First observed
edit_family_chat_message - First observed
export_accounts_csv - First observed
export_all_data - First observed
export_transactions_csv - First observed
forecast_cash_flow - First observed
get_account_details - First observed
get_account_health - First observed
get_account_history - First observed
get_agent_chat_session - First observed
get_agent_instructions - First observed
get_balance_sheet - First observed
get_budget_history - First observed
get_budget_streaks - First observed
get_cash_flow - First observed
get_category_suggestions - First observed
get_category_trend - First observed
get_celebrations - First observed
get_connection_health - First observed
get_daily_averages - First observed
get_debt_paydown - First observed
get_family_activity - First observed
get_family_group - First observed
get_financial_health_score - First observed
get_financial_ratios - First observed
get_forum_post - First observed
get_gamification_summary - First observed
get_goal_plan - First observed
get_goal_projection - First observed
get_income_statement - First observed
get_leaderboard - First observed
get_merchant_spending - First observed
get_monitoring_plan - First observed
get_net_worth - First observed
get_net_worth_history - First observed
get_period_summary - First observed
get_plaid_item - First observed
get_points_and_level - First observed
get_savings_rate_trend - First observed
get_scenario - First observed
get_seasons_summary - First observed
get_setup_status - First observed
get_shared_account_history - First observed
get_spending_anomalies - First observed
get_spending_summary - First observed
get_standing_facts - First observed
get_subscription_status - First observed
get_support_ticket - First observed
get_uncategorized_summary - First observed
get_unread_notification_count - First observed
get_user_preferences - First observed
get_user_profile - First observed
get_weekly_recaps - First observed
import_accounts - First observed
import_transactions - First observed
join_challenge - First observed
leave_challenge - First observed
leave_family_group - First observed
link_goal_plan - First observed
list_accounts - First observed
list_agent_chat_sessions - First observed
list_badges - First observed
list_budgets - First observed
list_categories - First observed
list_challenges - First observed
list_debt_payment_rules - First observed
list_debt_payments - First observed
list_family_chat_messages - First observed
list_family_invitations - First observed
list_family_members - First observed
list_forum_categories - First observed
list_forum_posts - First observed
list_goal_plans - First observed
list_goals - First observed
list_hard_assets - First observed
list_milestones - First observed
list_notifications - First observed
list_payment_history - First observed
list_payment_links - First observed
list_plaid_items - First observed
list_recurring_transactions - First observed
list_reminders - First observed
list_scenarios - First observed
list_shared_accounts - First observed
list_shared_goals - First observed
list_subscription_plans - First observed
list_support_tickets - First observed
list_tags - First observed
list_themes - First observed
list_transactions - First observed
mark_notification_read - First observed
open_billing_portal - First observed
override_account_balance - First observed
parse_statement - First observed
reactivate_subscription - First observed
recategorize_similar_transactions - First observed
reject_debt_payment_match - First observed
remove_family_member - First observed
rename_agent_chat_session - First observed
reorder_accounts - First observed
reply_to_support_ticket - First observed
reset_onboarding - First observed
revoke_family_invitation - First observed
save_payment_link - First observed
send_family_chat_message - First observed
send_family_invitation - First observed
start_checkout - First observed
suggest_debt_payment_pattern - First observed
switch_account - First observed
sync_plaid_item - First observed
undo_transaction - First observed
unlink_plaid_item - First observed
update_account - First observed
update_budget - First observed
update_category - First observed
update_debt_payment_rule - First observed
update_family_shares - First observed
update_forum_post - First observed
update_forum_reply - First observed
update_goal - First observed
update_hard_asset - First observed
update_milestone - First observed
update_monitoring_plan - First observed
update_recurring_transaction - First observed
update_scenario - First observed
update_tag - First observed
update_transaction - First observed
update_user_preferences - First observed
update_user_profile
Publisher details
- Operator
- Zoninga · Publisher source
- Operator website
- https://zoninga.com
- Vendor relationship
- First-party · Publisher source
- Documentation
- https://zoninga.com/ai/
- Trust center
- https://zoninga.com/security/
- Restrictions
- Requires a Zoninga account on the Premium plan; the 14-day free trial and comped accounts also qualify, but Free accounts cannot connect. Authorizing a connector additionally requires a verified email address and TOTP two-factor authentication on the Zoninga account. No custom OAuth app is needed: the endpoint supports RFC 7591 dynamic client registration, so the client registers itself. Bank linking covers US institutions through Plaid; manual entry and statement import work anywhere. On the client side, Claude custom connectors require a paid Claude plan, and ChatGPT connectors require Pro, Team, Enterprise, or Edu. · Publisher source
Related MCP Connectors
Personal-finance workspace for AI agents: accounts, spending, budgets, goals, and investments.
Personal finance for AI agents — onboard, import statements, categorize & budget over MCP.
Personal finance ledger for AI agents — query spending, track bills, forecast cash flow.
Track expenses, budgets, balances, transfers, and multi-currency reports with OAuth-secured tools.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceEnables AI assistants to manage personal finances through plain-language commands, including expense logging, account and deposit tracking, budget monitoring, and spending reports, with OAuth-secured access to a shared ledger.MIT
- AlicenseCqualityAmaintenanceopen-source personal finance app with a first-party MCP server. 91 HTTP tools (OAuth 2.1 + DCR) and 87 stdio tools cover transactions, budgets, accounts, portfolio analytics, FX conversion, loans, subscriptions, goals, importers, and rules. Users self-host with Docker + PostgreSQL or use the managed cloud8920AGPL 3.0
- AlicenseNot gradedqualityDmaintenanceEnables AI assistants to access and manage personal financial data from US institutions and manual entries, including transactions, balances, liabilities, and investments.MIT
- AlicenseNot gradedqualityBmaintenanceEnables compatible AI clients to read stored personal-finance data and, with explicit user authorization, perform supported mutations under the same ownership checks as the web app.MIT
Glama MCP Gateway
Add one secure layer between your agents and this server.