arcade-ynab-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@arcade-ynab-mcpshow me my budget for this month"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
arcade-ynab-mcp
An MCP server for YNAB, built with
arcade-mcp, hosted on Arcade Cloud via
arcade deploy, and used through Arcade MCP Gateways.
Each user connects their own YNAB account through OAuth 2.0 the first time they call a tool. The server holds no secrets. See docs/SPEC.md for the full design.
Tools
Phase 1 (read-only):
Tool | What it does |
| The authenticated YNAB user |
| The user's plans (budgets) |
| A plan's currency and date formats |
| Ready to Assign, age of money, and every category's amounts for a month |
| Monthly summaries, newest first |
| Accounts and balances |
| Categories, amounts and goals |
| Payees, optionally filtered by name |
| Transactions by date, account, category or payee |
| Upcoming and recurring transactions |
| Money moved between categories |
Phase 2 (writes):
Tool | What it does |
| Add a purchase, income, transfer or split |
| Approve, recategorize, flag or edit up to 100 transactions at once |
| Delete one transaction (destructive) |
| Import new transactions from linked accounts |
| Set a category's assigned amount for a month |
| Move money between categories or Ready to Assign |
| Create or edit categories and their targets |
| Create or rename category groups |
| Create or rename payees |
| Create an unlinked account |
| Manage scheduled transactions |
Phase 3 (summaries, read-only):
Tool | What it does |
| Unapproved transactions with suggested categories |
| Overspent categories and where to cover them from |
| Spending by category, category group or payee for a date range |
| Projected balances from scheduled transactions |
| Underfunded targets and what it takes to fund them |
Every tool is tagged read-only or write (and delete tools as destructive), so a gateway can expose only the read tools.
Amounts are in currency units (not YNAB milliunits). Every tool defaults to the user's most recently used plan.
Related MCP server: YNAB MCP Server
Setup
One-time setup of the YNAB OAuth app and the Arcade OAuth provider (ID ynab) is
described in docs/SPEC.md.
Development
uv tool install arcade-mcp # Arcade CLI
uv sync --extra dev # project and dev dependencies
uv run pytest # unit tests (YNAB is mocked; no network)
uv run ruff check . && uv run ruff format --check . && uv run mypy srcTool-selection evals (need an LLM API key; they don't call YNAB):
ANTHROPIC_API_KEY=... uv run arcade evals evals/ -p anthropicTo try the tools against your own YNAB account locally, log in with arcade login
(the OAuth flow runs through your Arcade project), then:
uv run src/arcade_ynab/server.py # stdio
uv run src/arcade_ynab/server.py http # streamable HTTP on :8000Deploy
arcade login
arcade deploy -e src/arcade_ynab/server.pyThen add the server's tools to an MCP Gateway in the Arcade dashboard.
Security
Never commit tokens or personal budget data. Tests use invented data only. YNAB tokens are held by Arcade and injected per request; they are never exposed to the model.
License
Available Tools
35 toolsYnab_AssignToCategoryAssignToCategoryAIdempotent
Set how much is assigned to a category for a month. The difference comes from, or goes back to, Ready to Assign.
| Name | Required | Description | Default |
|---|---|---|---|
| month | No | The plan month: 'current', or a date in the month such as '2026-03' or '2026-03-01'. | |
| amount | Yes | The new TOTAL assigned amount for the month in currency units (not an increment). Use MoveMoney to move a specific amount between categories. | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| category_id | Yes | The category to assign money to. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-read-only, idempotent, non-destructive behavior, so the bar is lower. The description adds genuinely useful behavioral context beyond them: the delta is drawn from or returned to Ready to Assign, which tells the agent the side effect is a reallocation of the plan's unassigned pool rather than a loss of funds.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, with the core action stated first and the funds-flow consequence second. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. For a money-mutation tool, the description supplies the key semantic (Ready to Assign as the counterparty), though it omits any note on permissions or overwrite behavior of the prior assignment.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description itself adds no parameter-level detail (no mention of month format, plan_id defaulting, or the total-vs-increment distinction), all of which is already handled by the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Set how much is assigned to a category for a month'), which is unambiguous. It does not, however, differentiate itself from the sibling Ynab_MoveMoney in the description text itself — that differentiation lives only in the parameter schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit when-to-use or when-not-to-use guidance; routing to MoveMoney for incremental moves appears in the `amount` schema description rather than the tool description. Usage is only implied by the 'assigned amount for a month' framing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_CreateAccountCreateAccountA
Create an unlinked account with a starting balance. Checking, savings, cash and credit card accounts are on budget; other assets and liabilities are tracking accounts.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The account name. | |
| balance | Yes | Starting balance in currency units. Use a negative number for money owed (credit cards and liabilities). | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| account_type | Yes | The kind of account. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is a non-read-only, non-idempotent, non-destructive, open-world write. The description adds genuinely useful behavioral context beyond that: the account is created *unlinked* (no bank connection), and the account_type choice determines whether the account counts toward the budget. It does not mention permission requirements or duplicate-name behavior, so not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the core action is front-loaded before the account-type consequence. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and annotations carry the safety profile. The description covers the mutation's scope and the account-type semantics; only minor gaps remain (permissions, duplicate handling) for a 4-param creation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; however the description adds semantic meaning the enum alone cannot convey — which account types are on-budget (checking, savings, cash, creditCard) versus tracking (otherAsset, otherLiability). That is real added value over the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create an unlinked account') plus the required starting balance, and clarifies the on-budget vs tracking distinction. It does not explicitly differentiate itself from any sibling (e.g., Ynab_ImportTransactions, which can also materialize accounts), so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case (opening a new manual/unlinked account) and explains which account_type values map to on-budget vs tracking accounts, but it never states when to use this versus alternatives or any prerequisite/ordering constraints. Implied usage only.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_CreateCategoryCreateCategoryC
Create a category in a category group, optionally with a target (goal).
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The new category's name. | |
| note | No | Optional note. | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| goal_target | No | Target amount in currency units. Creates a 'Needed for spending' target if the category has none. | |
| goal_frequency | No | Make the target repeat monthly, weekly or yearly. Requires goal_target. | |
| goal_target_date | No | Target date (YYYY-MM-DD). Can't combine with frequency. | |
| category_group_id | Yes | The category group to put it in. Use ListCategories to find group IDs. | |
| goal_needs_whole_amount | No | true = 'Set aside another' the target each period; false = 'Refill up to' the target. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds nothing beyond that: it does not warn that repeated calls create duplicate categories despite idempotentHint=false, nor mention permissions, errors, or whether a goal can later be modified.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or repetition. It is efficient, though it is arguably too terse for an 8-parameter mutation tool where a clause on goal interaction would earn its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and the parameter schema is fully documented, so return values and argument semantics need not be repeated. What is missing is usage context: when creating a category is appropriate versus updating an existing one, and the duplicate-creation behavior implied by non-idempotence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across all 8 parameters, and the schema already explains goal_target creating a 'Needed for spending' target, the frequency/date mutual exclusion, and ID lookups. The description's phrase 'optionally with a target (goal)' is weaker than what the schema already provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource (create a category) plus its placement constraint (in a category group) and an optional capability (target/goal). It does not differentiate itself from siblings like CreateCategoryGroup or CreatePayee, but the action itself is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no alternatives named, no prerequisites or exclusions stated. The only routing help (use ListPlans / ListCategories to find IDs) lives in the schema descriptions, not the tool description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_CreateCategoryGroupCreateCategoryGroupC
Create a category group.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The new group's name (up to 50 characters). | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the operation is a non-readonly, non-idempotent, open-world write (readOnlyHint=false, idempotentHint=false). The description adds no behavioral context beyond this — no mention of side effects on existing groups, ordering, or what happens on duplicate names, so it contributes nothing the structured fields don't already say.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero waste, but it is arguably too thin for a mutation tool whose siblings are easy to confuse it with. It is concise without being informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and annotations plus full schema coverage carry the safety and parameter load. Still, for a create/write tool with many sibling create/update tools, the description lacks any disambiguating context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (name, plan_id with its 'last-used' default and ListPlans pointer) are fully documented in the schema. The description adds no parameter meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a clear verb+resource ('Create a category group'), so an agent knows exactly what the tool does. It does not, however, differentiate itself from close siblings like Ynab_CreateCategory or Ynab_UpdateCategoryGroup in terms of scope or behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus CreateCategory, UpdateCategoryGroup, or ListCategories, and no prerequisites or exclusions are given. The agent must infer usage entirely from the name and sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_CreatePayeeCreatePayeeA
Create a payee. (CreateTransaction can also create a payee from payee_name.)
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The payee's name. | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false and openWorldHint=true, so the safety and repeat-call profile is covered structurally. The description adds the cross-tool behavioral note that CreateTransaction can create a payee too, which is useful, but it does not disclose what happens on duplicate names, whether a plan must be writable, or what is returned. Given the annotation coverage, 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero waste, purpose front-loaded and the alternative-path note parenthesized so it does not obscure the primary statement.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the parameter surface is fully documented in the schema. However, for a creation tool the description is silent on duplicate-payee behavior and on the relationship to Ynab_UpdatePayee, leaving gaps an agent would want filled before calling.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (name, plan_id) are documented in the schema, including the 'last-used' default and the ListPlans pointer. The description adds nothing beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Create a payee.' It also flags the overlap with Ynab_CreateTransaction, which lets an agent distinguish this tool from the most plausible alternative. It stops short of distinguishing from Ynab_UpdatePayee or explaining scope relative to plan_id, but the core purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical hints at an alternative route (creating a payee implicitly via CreateTransaction's payee_name), which implies a when-you-might-not-need-this condition, but it is not framed as guidance. There is no explicit statement of when to call this versus Ynab_UpdatePayee or when to prefer the implicit path, nor any prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_CreateScheduledTransactionCreateScheduledTransactionC
Schedule an upcoming or recurring transaction (bill, paycheck, transfer).
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | Date of the next occurrence (YYYY-MM-DD), in the future and within 5 years. | |
| memo | No | Optional memo. | |
| amount | Yes | Amount in currency units. Negative for outflows, positive for inflows. | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| payee_id | No | An existing payee ID. | |
| frequency | Yes | How often it repeats. 'never' means it happens once. | |
| account_id | Yes | The account it will be entered in. | |
| flag_color | No | Optional flag color. | |
| payee_name | No | Payee name (matches or creates a payee). Ignored if payee_id is given. | |
| category_id | No | Category ID, or an empty string for none. Split scheduled transactions aren't supported. | |
| transfer_account_id | No | For a scheduled transfer, the OTHER account's ID. Don't combine with a payee. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, so safety and idempotency are covered structurally. The description adds nothing beyond that — no note on duplicate scheduled entries, required permissions, or validation limits (e.g., the 5-year date bound is only in the schema).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with zero waste and the resource front-loaded. It is efficient, though the brevity is closer to under-specification than to disciplined conciseness for an 11-parameter creation tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given full schema coverage and an output schema, the description need not explain parameters or return values, so the core calling information is available. However, for a mutation tool with four required fields and transfer/payee interaction rules, a sentence about distinguishing scheduled creation from immediate creation would round it out.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with 11 well-documented parameters including enum values, so the schema carries the full semantic load. The description contributes no parameter meaning of its own; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb+resource ('Schedule ... transaction') and names concrete use cases (bill, paycheck, transfer), which separates it from Ynab_CreateTransaction's one-off creation. It does not explicitly name that sibling, but the recurring/upcoming framing is unambiguous about what gets created.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no mention of when to prefer Ynab_CreateTransaction instead, and no stated prerequisites (e.g., needing an existing account or payee). A reader can infer intent from 'recurring', but the routing decision between the two creation tools is left entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_CreateTransactionCreateTransactionB
Create a transaction: a purchase, income, a transfer between accounts, or a split.
| Name | Required | Description | Default |
|---|---|---|---|
| date | Yes | Transaction date (YYYY-MM-DD). Future dates are not allowed. | |
| memo | No | Optional memo. | |
| amount | Yes | Amount in currency units. Negative for an outflow (spending), positive for inflow. | |
| cleared | No | Cleared status. | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| approved | No | Mark as approved. Defaults to true, since the user asked for it explicitly. | |
| payee_id | No | An existing payee ID. | |
| account_id | Yes | The account the transaction belongs to. | |
| flag_color | No | Optional flag color. | |
| payee_name | No | Payee name. Matches an existing payee by name or creates a new one. Ignored if payee_id is given. | |
| category_id | No | Category ID. Leave empty for transfers between budget accounts and splits. | |
| subtransactions | No | Split lines. Each has an amount (same sign convention) and usually a category_id. The lines must add up to the transaction amount. | |
| transfer_account_id | No | For a transfer, the OTHER account's ID. The amount is from this account's point of view (negative moves money out of account_id). Don't combine with payee_id/payee_name. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the operation is a non-read-only, non-idempotent, open-world mutation, so the safety profile is covered. The description adds the transaction-type taxonomy (purchase/income/transfer/split) but does not disclose the behavioral consequences of idempotentHint=false (repeated calls create duplicates) or the payee auto-creation side effect, which the schema mentions but the description omits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler that immediately conveys the operation and its variants. Nothing needs trimming or reordering.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the schema carries the parameter burden. Still, for a 13-parameter mutation involving transfers and splits, the description is thin on when to prefer it over CreateScheduledTransaction or ImportTransactions and on the duplicate-creation risk.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the schema itself richly documents every parameter, including transfer sign conventions and split requirements. The description adds no parameter-level meaning beyond what the schema already provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a transaction') and enumerates the four supported shapes (purchase, income, transfer, split), which helps an agent understand the scope. It does not, however, distinguish this from close siblings like Ynab_CreateScheduledTransaction or Ynab_ImportTransactions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives such as CreateScheduledTransaction (future-dated), ImportTransactions (bulk), or UpdateTransactions. Usage is only implied by the verb 'Create'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_DeleteScheduledTransactionDeleteScheduledTransactionADestructiveIdempotent
Permanently delete a scheduled transaction. Only call this when the user has clearly asked to delete this specific scheduled transaction.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| scheduled_transaction_id | Yes | The exact ID of the scheduled transaction. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, readOnlyHint=false and idempotentHint=true, so the safety profile is partly covered. The description still adds value beyond them: 'Permanently delete' confirms irreversibility in user-facing terms, and the confirmation guardrail tells the agent it must have explicit user intent. It does not describe behavior when the ID does not exist, which is the main remaining gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste, with the destructive nature and the invocation precondition front-loaded. Nothing needs to be read twice or skimmed past.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and annotations carry the destructive/idempotency profile. For a single-resource destructive delete the description is nearly complete; it only omits what happens on an invalid or already-deleted ID.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters are documented in the schema (plan_id defaulting to 'last-used', scheduled_transaction_id as the exact ID). The description adds no parameter-level meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (delete) and resource (scheduled transaction), and the 'scheduled' qualifier is exactly what separates it from the sibling Ynab_DeleteTransaction and Ynab_UpdateScheduledTransaction. An agent can route correctly without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit gating condition: 'Only call this when the user has clearly asked to delete this specific scheduled transaction.' This is a clear when-to-use rule and a guard against speculative deletion, though it names no alternative (e.g. Ynab_UpdateScheduledTransaction) for the case where the user wants to change rather than remove it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_DeleteTransactionDeleteTransactionADestructiveIdempotent
Permanently delete one transaction. Only call this when the user has clearly asked to delete this specific transaction.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| transaction_id | Yes | The exact ID of the transaction to delete. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered. The description adds meaningful context by emphasizing that deletion is permanent and by imposing an explicit user-intent guardrail, though it says nothing about permissions, side effects on linked data, or response behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler, and the core action plus the critical call condition are front-loaded. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter destructive mutation, the description covers purpose and the key invocation guardrail, while annotations supply the safety profile and the output schema covers return values. Nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already fully documented in the schema, including the plan_id default and transaction_id meaning. The description adds no parameter-level detail beyond that, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Permanently delete one transaction') and the scope ('one') rules out bulk operations. It does not explicitly name a sibling such as DeleteScheduledTransaction, so it is clear but not fully differentiated from the tool family.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear, explicit condition for invocation: 'Only call this when the user has clearly asked to delete this specific transaction.' This is strong safety guidance for a destructive tool. It does not name alternative tools for related intents (e.g., deleting a scheduled transaction), which would have made it a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_FindOverspendingFindOverspendingARead-onlyIdempotent
Find overspent categories (negative available) in a month and suggest where to cover them from: Ready to Assign first, then the categories with the most money available that aren't saving toward an underfunded target. Use MoveMoney to cover after the user picks.
| Name | Required | Description | Default |
|---|---|---|---|
| month | No | The plan month: 'current', or a date in the month such as '2026-03' or '2026-03-01'. | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| max_candidates | No | How many categories to suggest as sources (1-20). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/idempotentHint/destructiveHint, so safety is covered. Beyond that, the description discloses the actual ranking algorithm (Ready to Assign first, then the categories with the most available money that aren't saving toward an underfunded target) and that human choice precedes any mutation, which is real behavioral context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler: the action leads, the selection heuristic follows, and the hand-off to MoveMoney closes. Every clause earns its place and the most decision-relevant information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-shape explanation is unnecessary, and the description still covers purpose, algorithmic behavior, and the next tool to call. For a read-only 3-parameter diagnostic tool, nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so 'month', 'plan_id', and 'max_candidates' are already fully documented in the schema. The description only indirectly reinforces the month scope and adds no syntax, defaults, or constraint detail for max_candidates. Baseline 3 applies when the schema carries the parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Find') and resource ('overspent categories (negative available)') with the scope ('in a month'), immediately distinguishing it from sibling query tools. It also names the sibling that consumes its output (MoveMoney), so the agent can place it precisely among 34 siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use and explicitly routes the follow-up action: 'Use MoveMoney to cover after the user picks,' which separates the read/diagnose step from the write step. It stops short of stating when-not to use it or any prerequisite (e.g., that the plan must be openable), but the workflow guidance is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_ForecastCashFlowForecastCashFlowARead-onlyIdempotent
Project account balances over the next N days from current balances and scheduled transactions (bills, paychecks, transfers). Reports each account's ending balance and its lowest point, to spot upcoming shortfalls. Doesn't include unscheduled spending.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | How many days ahead to forecast (1-365). | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| account_id | No | Only forecast this account. Defaults to all open budget accounts. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive, so the safety profile is covered. The description adds meaningful behavioral context beyond that: the forecast derives from current balances plus scheduled transactions only, and explicitly omits unscheduled spending, which is exactly the kind of limitation an agent needs to know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the core action and scope, followed by the output summary and the key limitation. No filler and no repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be re-explained, and annotations cover the safety profile. The description supplies the forecast basis and its exclusion, which is everything an agent needs to decide and call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters (days, plan_id, account_id) are already fully documented. The description's "next N days" phrasing adds nothing beyond the schema's "1-365" range. Baseline 3 is appropriate when structured fields carry the semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ("Project") with a clear resource ("account balances") and states the mechanism (current balances + scheduled transactions). It is distinguishable from neighbors like Ynab_FindOverspending or Ynab_SummarizeSpending, which look backward at past spending rather than projecting forward.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The scope note "Doesn't include unscheduled spending" implies when the tool is appropriate, but no sibling tool is named and no explicit when-to-use/when-not guidance is given. Usage must be inferred from the description's framing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_GetAccountGetAccountBRead-onlyIdempotent
Get a single account by ID.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| account_id | Yes | The account ID. Use ListAccounts to find it. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds nothing beyond that—no note on what is returned, whether a missing account errors, or plan scoping—so it contributes almost no behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the key qualifier ('single') front-loaded and zero waste. It borders on under-specification for a tool with a plan_id default, but nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only getter with a complete input schema, full annotations and an output schema, the description is nearly sufficient—return values and plan defaulting are handled elsewhere. Only the lack of any when-to-use routing keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and both parameters (account_id, plan_id with its 'last-used' default and ListPlans reference) are fully documented in the schema. The description adds no meaning beyond 'by ID', so the baseline 3 for schema-covered parameters applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Get') and resource ('a single account') and the lookup key (by ID), which cleanly separates it from Ynab_ListAccounts. It does not name the sibling, but the singular 'single account' implicitly contrasts with the list tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this instead of Ynab_ListAccounts or the other Get* siblings; the only routing hint ('Use ListAccounts to find it') lives in the schema, not the description. The agent must infer usage from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_GetCategoryGetCategoryARead-onlyIdempotent
Get a single category's amounts and goal progress for a given month.
| Name | Required | Description | Default |
|---|---|---|---|
| month | No | The plan month: 'current', or a date in the month such as '2026-03' or '2026-03-01'. | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| category_id | Yes | The category ID. Use ListCategories to find it. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered externally. The description adds only that the payload covers amounts and goal progress for a month, with no note on permission requirements, plan scoping, or error behavior for a missing category. With annotations carrying the burden, this is adequate but thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single 11-word sentence with the verb and scope front-loaded and zero filler. Nothing in it is redundant against the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and annotations cover the safety profile. For a simple single-record read, the description plus structured fields give an agent what it needs; only the lack of routing guidance keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are fully documented in the schema, making the baseline 3. The phrase 'for a given month' lightly reinforces the month parameter but adds no syntax or format detail beyond what the schema already provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Get'), resource ('category'), and scope ('amounts and goal progress for a given month'), which separates it from its ListCategories sibling at the read-single level. It does not explicitly name which sibling to use instead, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: the singular 'a single category' suggests this is the per-category lookup versus ListCategories for enumeration, but no when-to-use, when-not-to-use, or alternative routing is stated. There is no guidance on what to do when the category ID is unknown beyond the schema's pointer to ListCategories.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_GetMonthGetMonthARead-onlyIdempotent
Get one plan month: Ready to Assign, age of money, income, total assigned and activity, plus every category's assigned, activity and available amounts and goal progress.
Month totals cover all categories. When hidden categories are left out of the list, their combined amounts are in omitted_hidden_categories. Categories marked internal are YNAB system categories (e.g. Inflow: Ready to Assign), not budget lines.
This is the best single call for "how is my budget doing this month?".
| Name | Required | Description | Default |
|---|---|---|---|
| month | No | The plan month: 'current', or a date in the month such as '2026-03' or '2026-03-01'. | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| include_hidden | No | Include hidden categories. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, idempotent, non-destructive, openWorld), so the bar is lower, yet the description adds real nuance: hidden categories are excluded from the list and their combined amounts land in omitted_hidden_categories, and internal categories are system rows, not budget lines. That is genuinely useful context beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short paragraphs, front-loaded with the field inventory and followed by the caveats an agent needs. Every sentence carries information; the return-field list is dense but the useful caveats come before the routing line.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations, a full output schema and 100% schema coverage, the description only needs to frame scope and edge cases, which it does via the hidden/internal-category notes. It is complete for correct invocation, though the field enumeration partly restates what the output schema already provides.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so month, plan_id (including the 'last-used' default and ListPlans pointer) and include_hidden are already fully documented. The description alludes to hidden-category handling but adds no parameter syntax or format detail beyond the schema, matching the baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ('Get one plan month') and enumerates exactly what it returns: Ready to Assign, age of money, income, totals, and per-category assigned/activity/available plus goal progress. This clearly separates it from the plural sibling Ynab_ListMonths without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The closing line gives an explicit usage cue ('best single call for "how is my budget doing this month?"'), which tells the agent when this beats narrower calls like GetCategory or ListMonths. It stops short of naming alternatives or stating exclusions, so it is clear context rather than full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_GetPlanSettingsGetPlanSettingsBRead-onlyIdempotent
Get a plan's currency and date format settings.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false, and openWorldHint, and an output schema exists, so the safety and return profile are covered elsewhere. The description adds nothing beyond restating the returned fields — no mention of auth scope, rate limits, or what happens when the plan has no settings.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with zero filler, front-loaded with the verb and resource. Nothing could be trimmed without losing content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations covering the safety profile, an output schema covering return values, and a fully documented parameter, the description only needs to state the tool's scope — which it does. It is complete enough to call correctly, though a one-clause note on when to prefer it over ListPlans would close the remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single plan_id parameter is fully documented in the schema, including the 'last-used' default and the ListPlans pointer. Baseline 3 is appropriate since the description contributes no additional parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and resource ('a plan's currency and date format settings'), which is more informative than a bare name restatement. It does not, however, distinguish itself from sibling getters like Ynab_GetUser or Ynab_ListPlans beyond the resource noun.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no when-to-use guidance, no indication of when this is preferable to ListPlans or GetUser, and no prerequisites. The only routing hint ('Use ListPlans to find other plan IDs') lives in the schema, not the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_GetTransactionGetTransactionBRead-onlyIdempotent
Get a single transaction by ID, including its subtransactions if it is a split.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| transaction_id | Yes | The transaction ID. Use ListTransactions to find it. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety profile is fully covered. The description adds one useful behavioral fact — that split transactions return their subtransactions — which is real context beyond the annotations. It does not describe errors or missing-ID behavior, so a 3 is fitting.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no filler, front-loading the core action and immediately qualifying the split-transaction behavior. Nothing could be cut without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With rich annotations, 100% schema coverage, and an output schema to define the return shape, the description needs only to state the action and any non-obvious behavior — which it does. Only the lack of any alternative-tool routing keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already fully documented in the schema, including the plan_id default and ListPlans hint. The description's 'by ID' only loosely restates the required parameter and adds no format or constraint detail. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get a single transaction by ID') and adds a meaningful scope detail about subtransactions for split transactions. It is clearly distinct from list-style siblings by virtue of 'by ID', though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no when-to-use guidance, no prerequisites, and no mention of alternatives such as Ynab_ListTransactions. The schema text ('Use ListTransactions to find it') carries routing hints, but the description itself provides none.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_GetUserGetUserARead-onlyIdempotent
Get the YNAB user the current authorization belongs to.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds one genuinely useful behavioral fact beyond them — the returned user is derived from the current authorization rather than a supplied identifier — but says nothing about auth failure modes or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the scope qualifier follows the verb+resource immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read tool with a full annotation set and an output schema, the description carries exactly the burden it needs to — what is fetched and for whom. Return-value details are appropriately delegated to the output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so per the rubric the baseline is 4. The description correctly implies no input is needed, and there is nothing further for it to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ("Get") and a specific resource ("the YNAB user"), and the qualifier "the current authorization belongs to" pins down which user. No sibling tool (accounts, categories, plans, transactions, payees) returns the user, so there is no ambiguity to resolve.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: an agent can infer you call this to identify the authenticated user, but the description never states when to reach for it (e.g., before resolving budgets/plans) or that no alternative exists. No exclusions or prerequisites are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_ImportTransactionsImportTransactionsA
Ask YNAB to import new transactions from the plan's linked (direct import) accounts, like pressing "Import" in the app. Imported transactions arrive unapproved.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the safety profile is covered. The description adds genuine beyond-schema context with 'Imported transactions arrive unapproved,' telling the agent the resulting state of the imported records and implying a follow-up ReviewUnapproved step. It does not discuss delays, failure modes when no linked accounts exist, or what the response contains, so it is helpful rather than exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with the core action and its scope. Every clause carries information: the resource, the account restriction, the UX analog, and the resulting transaction state.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the single parameter is fully documented. For a non-read-only, open-world operation the description covers what it does and the state of the results, but it omits any note on prerequisites (linked accounts/direct import setup) or expected latency that an agent invoking a bank-sync operation might want.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (plan_id) with 100% schema description coverage, including the 'last-used' default and a pointer to ListPlans. The description adds nothing about plan_id, so the baseline 3 applies since the schema does all the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (import) plus resource (transactions) and narrows scope to the plan's linked direct-import accounts, which cleanly separates it from siblings like Ynab_CreateTransaction and Ynab_ListTransactions. The 'like pressing Import in the app' analogy makes the operation unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The app-analogy gives clear context for when this is the right call (pulling fresh transactions from bank-linked accounts rather than entering them manually), which differentiates it from CreateTransaction. It stops short of explicit when-not-to-use guidance or naming alternatives, and never notes that the plan must have direct import configured.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_ListAccountsListAccountsBRead-onlyIdempotent
List a plan's accounts with their type, whether they're on budget (or tracking only), and their working, cleared and uncleared balances.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| include_closed | No | Include closed accounts. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is fully covered. The description adds the informational content of the response (budget status and balance breakdowns), which is useful but partly duplicated by the output schema. No pagination, ordering, or closed-account-default behavior is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence that front-loads the action and resource, then enumerates the returned fields. No wasted wording, though the enumeration is dense and could be trimmed since the output schema covers it.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a full output schema, rich annotations, and 100% schema coverage, the description only needs to establish scope, which it does. The one gap is that it doesn't state whether closed accounts are returned by default, leaving the include_closed default ambiguous without the schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (plan_id with its 'last-used' default and ListPlans pointer, include_closed) are fully documented in the schema. The description adds nothing about parameter syntax or the default for include_closed, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('List') and resource ('a plan's accounts') and enumerates the returned content: type, on-budget vs tracking status, and the three balance figures. It is clearly distinguishable from Ynab_GetAccount and Ynab_CreateAccount by the plural 'list' framing, though it never explicitly names a sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, no prerequisites, and no routing to alternatives. It does not say to use Ynab_GetAccount for a single account, nor does it explain when to set include_closed. Usage is left entirely to inference from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_ListCategoriesListCategoriesARead-onlyIdempotent
List a plan's category groups and categories, with each category's assigned, activity and available amounts for the current month and its goal, if any.
Use GetMonth for amounts in a different month.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| include_hidden | No | Include hidden category groups and categories. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is fully covered by structured data. The description adds one real behavioral trait beyond that: the amounts returned are scoped to the current month and goals are included only 'if any'. No pagination, ordering, or hidden-category default behavior is described, though the hidden flag is documented in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler. The capability and its return shape are front-loaded, and the sibling redirect is a compact trailing sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be enumerated, and the description still summarizes them usefully. For a read-only, zero-required-parameter list tool with full schema coverage, this is essentially complete; only minor scope questions (ordering, pagination) remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters, including the 'last-used' default and the ListPlans cross-reference, so the schema does the heavy lifting. The description adds no parameter-level detail beyond implying the current-month scope; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) plus the exact resources (category groups and categories) and even the payload shape: assigned, activity and available amounts for the current month plus goal. This distinguishes it from Ynab_GetCategory (single category) and Ynab_GetMonth (monthly aggregates) without the agent opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes one alternative case: 'Use GetMonth for amounts in a different month.' That is a genuine when-to-use rule. It does not, however, address when to prefer this over GetCategory or how it relates to ListMonths, so it stops short of full alternative coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_ListMoneyMovementsListMoneyMovementsARead-onlyIdempotent
List money moved between categories, or between a category and Ready to Assign, newest first. Movements made together in one action share a money_movement_group_id.
A missing from_category_id means the money came from Ready to Assign; a missing to_category_id means it went to Ready to Assign. Use ListCategories to resolve IDs.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return (1-500). | |
| month | No | Only include movements in this plan month: 'current', '2026-03' or '2026-03-01'. Omit for all months. | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive and open-world, so the safety profile is covered. The description adds genuinely non-obvious behavior: result ordering (newest first), grouping of same-action movements via money_movement_group_id, and the null-category convention meaning Ready to Assign.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core purpose, then domain conventions, then a routing hint. Nothing is redundant, though the middle sentence on null IDs is dense for a description rather than schema documentation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the description covers the main interpretive quirks an agent needs (ordering, grouping, Ready to Assign convention). Adequate for a zero-required-param read tool; only minor gaps remain, such as default page size behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the three input parameters (limit, month, plan_id), so the schema carries the load and baseline 3 applies. The description's notes about missing from_category_id/to_category_id explain response-field semantics rather than the documented inputs, so it does not add input-level parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List money moved between categories, or between a category and Ready to Assign') plus ordering ('newest first'), which cleanly separates it from ListTransactions and MoveMoney. An agent can identify the tool without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context (listing historical movements) and a concrete prerequisite routing hint: 'Use ListCategories to resolve IDs.' It does not explicitly contrast with MoveMoney or ListTransactions, but the read-vs-write distinction is implied and the ID-resolution guidance is actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_ListMonthsListMonthsARead-onlyIdempotent
List monthly summaries for a plan (Ready to Assign, income, assigned, activity, age of money), newest first. Use GetMonth for category detail in a single month.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of months to return, newest first (1-500). | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered by structured data. The description adds useful behavior only in the ordering rule ('newest first'); it says nothing about pagination or how the limit interacts with the plan scope.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler. The core purpose and returned fields are front-loaded, and the sibling routing is a compact second sentence.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the schema handles defaults and limits. For a read-only list tool with a sibling named for the alternate path, nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (limit range 1-500, plan_id defaulting to 'last-used') are already fully documented. The description's 'newest first' merely echoes the schema's ordering note, adding no meaning beyond structured data. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (list monthly summaries for a plan) and enumerates the returned fields (Ready to Assign, income, assigned, activity, age of money). It explicitly distinguishes itself from the sibling GetMonth, so an agent can route without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the alternative tool and the condition that selects it: 'Use GetMonth for category detail in a single month.' This gives clear routing guidance, though it does not state exclusions in the other direction or mention the ListPlans prerequisite for non-default plans (the schema does, however).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_ListPayeesListPayeesARead-onlyIdempotent
List a plan's payees, sorted by name. Transfer payees (which represent transfers to another account) include the transfer_account_id.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return (1-500). | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| name_contains | No | Only return payees whose name contains this text (case-insensitive). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds real behavioral value beyond that: results are sorted by name, and transfer payees carry a transfer_account_id, which tells the agent how to interpret certain rows. Pagination/limit behavior is left to the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action and followed by the one nuance worth knowing. No filler and nothing repeated from the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation, and annotations cover the safety profile. The description supplies the remaining decision-relevant nuance (sort order, transfer payee shape) so an agent can call and interpret this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with limit, plan_id, and name_contains each documented including the 'last-used' default and the 1-500 range. The description adds no parameter-level detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List a plan's payees') and immediately narrows scope with 'sorted by name'. The resource itself (payees) cleanly separates it from siblings like Ynab_ListAccounts, Ynab_ListCategories, and Ynab_ListTransactions, and the note about transfer payees adds a distinguishing detail.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: an agent can infer this is the tool for enumerating payees when it needs payee IDs. There is no explicit when-to-use vs. alternatives guidance and no mention that filtering is available via the schema's name_contains, though the plan_id field points at Ynab_ListPlans.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_ListPlansListPlansBRead-onlyIdempotent
List the user's YNAB plans (budgets), most recently modified first.
| Name | Required | Description | Default |
|---|---|---|---|
| include_accounts | No | Also return each plan's open and closed accounts. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds one genuinely useful behavioral detail not in the annotations — result ordering 'most recently modified first' — but discloses nothing about pagination or result-size limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the resource and the ordering constraint are both stated immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be described, and annotations cover safety, so the description is nearly sufficient for a zero-required-parameter list call. The only real omission is pagination/result-size behavior, which a 'list all plans' tool might reasonably need.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single parameter include_accounts, so the schema already documents it fully. The description never mentions the parameter or its expansion behavior, so it adds no meaning beyond structured data — the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List the user's YNAB plans') and disambiguates YNAB terminology by equating plans with budgets, which helps separate it from siblings like ListAccounts or ListMonths. It does not explicitly name a sibling it is not, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use, prerequisites, or alternative-routing guidance (e.g., use GetPlanSettings for a single plan's configuration). Usage is only inferable from the verb, which is weak for a tool sharing a namespace with 30+ siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_ListScheduledTransactionsListScheduledTransactionsARead-onlyIdempotent
List upcoming and recurring scheduled transactions, soonest first, with their frequency and next date. Amounts are negative for outflows, positive for inflows.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return (1-500). | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so the bar is lower, and the description still adds real behavioral context: results are sorted soonest-first, entries carry frequency and next date, and amounts use a signed convention (negative outflow, positive inflow). It does not mention pagination or total counts, which is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, zero waste; the ordering and sign conventions an agent most needs are front-loaded. Nothing is redundant with the annotations or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return-value detail is unnecessary, and annotations cover the safety profile; the description fills in ordering and the amount sign convention. Slightly short on pagination/limit behavior given the 1-500 cap, but essentially complete for this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (limit, plan_id) are fully documented in the schema, including the 'last-used' default and the ListPlans pointer. The description adds nothing about parameters, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('List ... scheduled transactions') and narrows scope to 'upcoming and recurring', which implicitly separates it from Ynab_ListTransactions. It does not explicitly name the sibling it differs from, but an agent can tell what it returns without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the resource name (scheduled vs. actual transactions) but there is no explicit when-to-use, when-not-to-use, or pointer to alternatives such as Ynab_ListTransactions or Ynab_CreateScheduledTransaction. Minimum-viable routing only.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_ListTransactionsListTransactionsARead-onlyIdempotent
List transactions, newest first, optionally filtered by date range and by ONE of account, category or payee.
Amounts are in currency units: negative is an outflow, positive is an inflow. Split transactions include their subtransactions. When filtering by category or payee, split lines are returned as their own rows with type 'subtransaction'.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return (1-500). | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| payee_id | No | Only include transactions with this payee. | |
| account_id | No | Only include transactions in this account. | |
| since_date | No | Only include transactions on or after this date (YYYY-MM-DD). YNAB defaults to one year ago when omitted. | |
| until_date | No | Only include transactions on or before this date (YYYY-MM-DD). | |
| category_id | No | Only include transactions in this category. | |
| transaction_type | No | Only include unapproved or uncategorized transactions. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the annotations (which only cover readOnly/idempotent/non-destructive): the newest-first ordering, the sign convention for amounts (negative outflow, positive inflow), and the non-obvious behavior that split lines surface as separate rows typed 'subtransaction' when filtering by category or payee. These are exactly the traits an agent cannot infer from structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each load-bearing: scope/ordering, filter exclusivity, amount sign convention, split-row behavior. The most decision-relevant information (what it lists, how to filter) is front-loaded with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With read-only annotations and an output schema present, the description need not explain return values or safety. It covers ordering, filter rules, sign semantics, and split behavior — everything an agent needs to call and interpret this correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds a real constraint absent from the schema: account_id, category_id and payee_id are mutually exclusive ('ONE of'). That mutual-exclusivity rule is genuine parameter meaning beyond the per-parameter schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('List transactions') and immediately scopes it with ordering ('newest first') and filter conditions. The 'ONE of account, category or payee' clause distinguishes this bulk-list tool from the single-item sibling Ynab_GetTransaction. An agent can tell exactly what it gets back without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States a concrete usage rule — filters are mutually exclusive across account/category/payee — which guides invocation. It does not name alternative tools such as Ynab_GetTransaction or Ynab_ListScheduledTransactions, nor state when-not to use it, so it stops short of explicit routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_MoveMoneyMoveMoneyA
Move money between two categories, or between a category and Ready to Assign, by adjusting their assigned amounts for the month. Running it twice moves the money twice.
The source must have at least the amount available. YNAB has no single "move" call, so this updates the source first and then the destination. If the second step fails, the money is left in Ready to Assign and the error says so.
| Name | Required | Description | Default |
|---|---|---|---|
| month | No | The plan month: 'current', or a date in the month such as '2026-03' or '2026-03-01'. | |
| amount | Yes | Amount to move, in currency units (positive). | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| to_category_id | No | Category to give money to. Omit to return it to Ready to Assign. | |
| from_category_id | No | Category to take money from. Omit to take it from Ready to Assign. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: it warns that repeated invocation moves money repeatedly (reinforcing idempotentHint=false), explains the two-step source-then-destination implementation YNAB forces, and spells out the partial-failure outcome where funds land in Ready to Assign with a matching error. This is exactly the behavioral detail an agent needs before mutating budgets.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short, front-loaded sentences with no filler: purpose first, precondition second, failure semantics last. Every sentence carries information an agent would otherwise have to discover empirically.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a non-idempotent mutation tool with an output schema and full parameter coverage, the description supplies what the structured fields cannot: the side-effect cadence, the two-call composition, and the degraded-failure state. Nothing essential to calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents month, amount, plan_id, and the from/to category IDs, including the 'omit to use Ready to Assign' semantics. The description confirms the source/destination direction of the move but adds no syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Move money between two categories, or between a category and Ready to Assign,' and clarifies the mechanism ('adjusting their assigned amounts'). It is clear on its own, but it never names the nearest sibling Ynab_AssignToCategory, which an agent could easily confuse with this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an implicit precondition ('The source must have at least the amount available') and describes the effect, but offers no explicit when-to-use guidance and no routing to or from the closely related AssignToCategory. Usage must be inferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_ReviewGoalsReviewGoalsARead-onlyIdempotent
Review category targets (goals) for a month: which are underfunded and by how much, how many are on track, and whether Ready to Assign can cover the shortfall. Use AssignToCategory or MoveMoney to fund them after the user confirms.
| Name | Required | Description | Default |
|---|---|---|---|
| month | No | The plan month: 'current', or a date in the month such as '2026-03' or '2026-03-01'. | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive and openWorld, so the safety profile is covered. The description adds genuine context beyond that: it discloses the analysis performed (underfunding amounts, on-track count, Ready to Assign coverage) and warns that any funding action needs user confirmation, which is workflow guidance annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler: the capability and its output facets come first, the next-step routing second. Nothing redundant with the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and a 100%-covered schema handles both parameters. Combined with annotations carrying the safety profile, the description supplies everything an agent needs to select and invoke it, including the downstream workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (month, plan_id) already document defaults and accepted formats, so the schema does the heavy lifting. The description only implies the month scoping and adds no syntax or default detail beyond it — baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource ('Review category targets (goals) for a month') followed by an explicit enumeration of what the review surfaces: underfunded categories and shortfall, count on track, and Ready to Assign coverage. This clearly separates it from GetMonth/ListCategories and from the mutation siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for when to run it and names the follow-up tools (AssignToCategory, MoveMoney) with a prerequisite ('after the user confirms'). There is no explicit when-not or 'use X instead' exclusion, so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_ReviewUnapprovedReviewUnapprovedARead-onlyIdempotent
List transactions waiting for approval, each with a suggested category based on how the same payee was categorized recently. Use UpdateTransactions to approve or recategorize them after the user confirms.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of results to return (1-500). | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| since_date | No | Only include unapproved transactions on or after this date (YYYY-MM-DD). Defaults to one year ago, YNAB's default. | |
| history_days | No | How many days of approved transactions to learn payee categories from (7-365). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, so safety is covered. The description adds genuinely non-obvious behavior beyond them: each result carries a *computed* suggested category learned from recent approved payee history, and it flags that a user confirmation step precedes any write.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core listing behavior, then the follow-up workflow. No filler and nothing redundant with the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and annotations cover the safety profile. What remains — the derived suggestion and the next action — is stated, making the definition complete for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents limit, plan_id, since_date and history_days. The description alludes to the history mechanism ('how the same payee was categorized recently'), which lightly maps to history_days, but adds no format or range detail. Baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (list) plus resource (transactions waiting for approval) plus a distinctive behavior (suggested category derived from recent payee history). This clearly separates it from the generic sibling Ynab_ListTransactions and from Ynab_GetTransaction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes the follow-up action to UpdateTransactions and conditions it on user confirmation, so the agent knows this tool is read-only reconnaissance. It does not explicitly state when to prefer this over ListTransactions, but the 'waiting for approval' scoping makes it inferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_SummarizeSpendingSummarizeSpendingARead-onlyIdempotent
Summarize spending over a date range, grouped by category, category group or payee.
Counts budget-account transactions only. Transfers between accounts and income (inflows to Ready to Assign) are excluded; refunds reduce spending in their category. Spending is reported as a positive number.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of groups to return, biggest first (1-500). | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| group_by | No | Group spending by category, category group or payee. | |
| since_date | Yes | Start of the period (YYYY-MM-DD), inclusive. | |
| until_date | No | End of the period (YYYY-MM-DD), inclusive. Defaults to today. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive behavior, so the safety bar is low. The description still adds genuine business semantics beyond the annotations: it counts budget-account transactions only, excludes transfers and income inflows, applies refunds as reductions, and normalizes spending to a positive number.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action, followed by the inclusion/exclusion rules and the sign convention. Every sentence carries distinct information with no waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not cover return values. It fully covers scope, filtering rules, and grouping semantics for a read-only summary tool, leaving no material gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters including the group_by enum are already documented. The description reinforces the grouping dimension but adds no format or default detail beyond what the schema provides, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (summarize) and resource (spending) with its scope (date range) and grouping options (category, category group, payee). An agent can distinguish it from analytical siblings like ForecastCashFlow or FindOverspending without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when it applies by defining what counts as spending, but it never explicitly names alternatives or states when to choose this over ForecastCashFlow or FindOverspending. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_UpdateCategoryUpdateCategoryAIdempotent
Rename a category, change its note or group, or set or remove its target (goal). Only the fields you pass change. (YNAB's API can't hide or unhide categories.)
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | New name. | |
| note | No | New note. Pass an empty string to clear it. | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| category_id | Yes | The category to change. | |
| goal_target | No | Target amount in currency units. Creates a 'Needed for spending' target if the category has none. | |
| remove_goal | No | Remove the category's target entirely. | |
| goal_frequency | No | Make the target repeat monthly, weekly or yearly. Requires goal_target. | |
| goal_target_date | No | Target date (YYYY-MM-DD). Can't combine with frequency. | |
| category_group_id | No | Move the category to this group. | |
| goal_needs_whole_amount | No | true = 'Set aside another' the target each period; false = 'Refill up to' the target. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-destructive, idempotent, open-world semantics, so the bar is lower. The description adds two genuinely useful behaviors beyond them: patch semantics ('Only the fields you pass change') and an API-level limitation ('YNAB's API can't hide or unhide categories') that prevents a failed or misdirected call.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both front-loaded and zero-waste: capabilities first, then the patch semantics and the hard limitation. Nothing redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema, full parameter documentation, and annotations covering safety, the description is nearly sufficient. It could still note permission requirements or what happens to unmentioned fields' side effects on the plan, but nothing an agent needs to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, including cross-field constraints ('Requires goal_target', 'Can't combine with frequency'), so the schema carries the load. The description restates the target/group/note fields at a high level without adding syntax or interaction rules beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (rename/change/set/remove) against a specific resource (category), and enumerates exactly the mutable surfaces: name, note, group, target/goal. This clearly distinguishes it from Ynab_CreateCategory and Ynab_UpdateCategoryGroup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied (modify an existing category) but never stated explicitly, and no when-not guidance or alternatives are named despite siblings like Ynab_UpdateCategoryGroup and Ynab_CreateCategory being easy to confuse with this tool. The partial-update note is behavior, not selection guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_UpdateCategoryGroupUpdateCategoryGroupBIdempotent
Rename a category group.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The new name (up to 50 characters). | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| category_group_id | Yes | The category group to rename. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false, covering the safety profile. 'Rename' adds one useful behavioral detail beyond the annotations: this mutation only changes the name field rather than arbitrary category-group properties. It still omits permissions, side effects, and error behavior, so it remains thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, which is structurally efficient. It is arguably too terse for a mutation with three parameters, but there is no wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, annotations covering safety/idempotency, and 100% schema description coverage, the core facts an agent needs are available. The description itself contributes almost nothing beyond the verb, which is minimally adequate for a low-complexity rename operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are already documented in the schema, including plan_id defaults and the 50-character name limit. The description adds no parameter-level detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Rename') and resource ('category group'), clearly separating it from siblings like Ynab_UpdateCategory and Ynab_CreateCategoryGroup. It does not explicitly name or contrast with those siblings, but the scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisite conditions, and no mention of alternatives such as Ynab_UpdateCategory. The agent gets no help distinguishing this from other update tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_UpdatePayeeUpdatePayeeBIdempotent
Rename a payee. The new name shows on all of its transactions.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | The new name. | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| payee_id | Yes | The payee to rename. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the mutation profile (readOnlyHint=false, destructiveHint=false, idempotentHint=true, openWorldHint=true), so the safety story is largely told. The description adds one genuinely useful non-obvious fact: the rename propagates to all of the payee's existing transactions. It does not mention permissions or error behavior, so it is moderate added value rather than rich.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero waste, and the core action is front-loaded before the side-effect note. Nothing could be cut without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no explanation, and annotations cover the safety/idempotency profile. The description supplies the one behavioral fact an agent needs (transaction propagation) and the schema covers defaults, leaving only minor gaps such as permissions or failure modes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so payee_id, name, and plan_id (including the 'last-used' default and the ListPlans pointer) are fully documented in the schema itself. The description adds no parameter-level detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource: 'Rename a payee.' This clearly separates it from sibling tools like Ynab_CreatePayee and Ynab_ListPayees by the distinct 'rename' verb. It stops short of explicitly naming or contrasting those siblings, so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no reference to alternatives (e.g., why rename via this tool rather than Ynab_UpdateTransactions). The agent is left to infer that this is the only path for renaming a payee.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_UpdateScheduledTransactionUpdateScheduledTransactionAIdempotent
Change a scheduled transaction. Only the fields you pass change.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | New next date (YYYY-MM-DD), in the future. | |
| memo | No | New memo. Pass an empty string to clear it. | |
| amount | No | New amount in currency units. | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| payee_id | No | An existing payee ID. | |
| frequency | No | New repeat frequency. | |
| account_id | No | Move it to this account. | |
| flag_color | No | New flag color, or 'none' to remove it. | |
| payee_name | No | Payee name (matches or creates a payee). Ignored if payee_id is given. | |
| category_id | No | Category ID, or an empty string for none. Split scheduled transactions aren't supported. | |
| transfer_account_id | No | For a scheduled transfer, the OTHER account's ID. Don't combine with a payee. | |
| scheduled_transaction_id | Yes | The scheduled transaction to change. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=false, idempotentHint=true and openWorldHint=true, so safety and repeatability are covered. The description adds a genuinely non-annotated behavior: partial-update semantics where omitted fields are preserved, which tells the agent it is not a full replace and that clearing values requires explicit empty strings (schema). No auth or error behavior is described.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, zero filler, and the patch semantics are front-loaded immediately after the purpose. Nothing could be trimmed without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the fully documented 12-parameter schema plus annotations carry the heavy lifting. The remaining gap is the absence of any when-to-use routing against the many sibling update tools, which for a mutation with 12 options would have been worth a sentence.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and every parameter carries its own description, including the payee_id-vs-payee_name precedence, the 'none' flag sentinel, and the split-transaction limitation. The description adds no parameter-level detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Change a scheduled transaction'), which cleanly separates it from the Create/Delete/List scheduled-transaction siblings. It stops short of naming the routing alternatives (e.g., CreateScheduledTransaction for new entries, UpdateTransactions for actuals), so it is clear but not fully differentiated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Only the fields you pass change' qualifies how to call it, but the description never says when to choose this over CreateScheduledTransaction, DeleteScheduledTransaction, or UpdateTransactions, nor does it state any prerequisite such as first resolving the scheduled_transaction_id. Usage must be inferred from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Ynab_UpdateTransactionsUpdateTransactionsAIdempotent
Apply the same changes to one or more transactions in a single request, e.g. approve a batch, recategorize, mark cleared, flag, or fix a memo. Only the fields you pass change.
| Name | Required | Description | Default |
|---|---|---|---|
| date | No | Change the date (YYYY-MM-DD). | |
| memo | No | Replace the memo. Pass an empty string to clear it. | |
| amount | No | Change the amount in currency units. Only allowed when updating a single transaction; split amounts can't be changed. | |
| cleared | No | Set the cleared status. | |
| plan_id | No | The YNAB plan (budget) ID. Defaults to 'last-used', the plan the user opened most recently. Use ListPlans to find other plan IDs. | |
| approved | No | Set approved (true) or unapproved (false). | |
| payee_id | No | Change the payee to this existing payee, or pass an empty string to remove the payee. | |
| flag_color | No | Set the flag color, or 'none' to remove it. | |
| payee_name | No | Change the payee by name (matches or creates a payee). | |
| category_id | No | Recategorize, or pass an empty string to make it uncategorized. Not allowed on split transactions. | |
| transaction_ids | Yes | IDs of the transactions to change (1-100). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and destructiveHint=false, and the description adds a genuinely important behavioral trait not present in structured fields: partial-update semantics ('Only the fields you pass change'). It omits auth/plan-scope caveats and error behavior on splits, leaving some room.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the operation and scope, followed by the partial-update rule. No padding or restated title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, and the batch limit and per-field constraints live in the schema. The description covers operation, scope, and partial-update behavior, but says nothing about failure modes (e.g., mixed split transactions in a batch) for an 11-param mutation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented — including split constraints on amount/category_id and the empty-string clearing conventions. The description's use-case list loosely maps to fields but adds no syntax or format detail beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb plus resource ('apply the same changes to one or more transactions in a single request') and immediately signals the batch/patch semantics. An agent can distinguish this from CreateTransaction, UpdateScheduledTransaction, and DeleteTransaction without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description enumerates concrete use cases — 'approve a batch, recategorize, mark cleared, flag, or fix a memo' — which tells the agent when this tool is the right choice. It stops short of naming alternatives (e.g., UpdateScheduledTransaction for scheduled transactions), so it does not meet the explicit when-not bar.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
35 tool updates
v0.3.0- First observed
Ynab_AssignToCategory - First observed
Ynab_CreateAccount - First observed
Ynab_CreateCategory - First observed
Ynab_CreateCategoryGroup - First observed
Ynab_CreatePayee - First observed
Ynab_CreateScheduledTransaction - First observed
Ynab_CreateTransaction - First observed
Ynab_DeleteScheduledTransaction - First observed
Ynab_DeleteTransaction - First observed
Ynab_FindOverspending - First observed
Ynab_ForecastCashFlow - First observed
Ynab_GetAccount - First observed
Ynab_GetCategory - First observed
Ynab_GetMonth - First observed
Ynab_GetPlanSettings - First observed
Ynab_GetTransaction - First observed
Ynab_GetUser - First observed
Ynab_ImportTransactions - First observed
Ynab_ListAccounts - First observed
Ynab_ListCategories - First observed
Ynab_ListMoneyMovements - First observed
Ynab_ListMonths - First observed
Ynab_ListPayees - First observed
Ynab_ListPlans - First observed
Ynab_ListScheduledTransactions - First observed
Ynab_ListTransactions - First observed
Ynab_MoveMoney - First observed
Ynab_ReviewGoals - First observed
Ynab_ReviewUnapproved - First observed
Ynab_SummarizeSpending - First observed
Ynab_UpdateCategory - First observed
Ynab_UpdateCategoryGroup - First observed
Ynab_UpdatePayee - First observed
Ynab_UpdateScheduledTransaction - First observed
Ynab_UpdateTransactions
TDQS
Scored across 35 tools
Tools mostly target distinct resources and actions, with clear list/get/create/update/delete boundaries. Minor overlap exists between AssignToCategory and MoveMoney, and between FindOverspending and ReviewGoals, but descriptions clarify the distinction.
Every tool uses the same Ynab_ prefix followed by PascalCase verb-object or verb-phrase naming. The pattern is predictable and consistent across all 35 tools.
35 tools is heavy for an MCP server, exceeding the 3-15 well-scoped range by a wide margin. While YNAB is a broad domain, the surface likely creates choice overload and could be consolidated.
The tool set covers YNAB's core lifecycle for transactions, scheduled transactions, accounts, categories, category groups, payees, months, plans, money movements, goals, approvals, spending summaries, forecasting, imports, user and settings. Missing singular getters are workarounds via list operations, and unsupported deletes align with API limits.
Maintenance
Related MCP Connectors
Read your accounts, budgets and net worth, and draft changes you confirm.
Personal finance by conversation: expenses, receipts, statement import, budgets, net worth.
Personal finance for AI agents: accounts, budgets, goals, 9-strategy debt payoff, reports. OAuth 2.1
Chat with your bank data: balances, transactions, budgets, bills. Reads only, never moves money.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables AI assistants to interact with YNAB budgets through natural language. Supports managing accounts, categories, transactions, and budget months with 21 tools for comprehensive budget operations.-
- AlicenseAqualityCmaintenanceEnables interaction with You Need A Budget (YNAB) through their API, allowing users to manage budgets, accounts, categories, transactions, payees, and scheduled transactions through natural language.1213 npm1GPL 3.0
- AlicenseAqualityBmaintenanceEnables AI assistants to interact with YNAB budgets, performing read-only queries by default and optional write operations like creating transactions and managing categories through natural language.39259 npm34MIT
- AlicenseAqualityDmaintenanceEnables reading and writing YNAB budget data, such as listing budgets, accounts, categories, transactions, and creating or updating transactions, through natural language commands.8MIT