Skip to main content
Glama
krax1337

mcp-server-zenmoney

by krax1337

mcp-server-zenmoney

CI npm

MCP server for ZenMoney (Дзен-мани). It gives an AI agent access to your accounts and balances, lets it search transactions and build reports, read and set monthly budgets, see planned (recurring) operations and debts, categorize spending, and record or edit transactions with safety rails. It talks to the official ZenMoney API v8 (POST /v8/diff/ sync protocol and POST /v8/suggest/) using an OAuth token, e.g. one issued through Zerro.

Unofficial project, not affiliated with ZenMoney or Zerro. Your token and data stay on your machine.

Quick start

Requirements: Node.js >= 22.

  1. Get a token: open https://zerro.app/token, log in with your ZenMoney account and copy the token shown.

  2. Verify it: ZENMONEY_TOKEN=... npx -y mcp-server-zenmoney --check prints OK: <login>, main currency …, N active accounts, M transactions (exit code 1 with the reason otherwise).

  3. Add the server to your agent (below).

Without a token the server still starts; every tool then explains how to configure one.

Related MCP server: Money Lover MCP Server

Connect to an agent (stdio)

Claude Code

claude mcp add zenmoney -e ZENMONEY_TOKEN=YOUR_TOKEN -- npx -y mcp-server-zenmoney

Claude Desktop / Cursor / most MCP clients

claude_desktop_config.json, ~/.cursor/mcp.json, …:

{
  "mcpServers": {
    "zenmoney": {
      "command": "npx",
      "args": ["-y", "mcp-server-zenmoney"],
      "env": { "ZENMONEY_TOKEN": "YOUR_TOKEN" }
    }
  }
}

omp

User-wide config ~/.omp/agent/mcp.json (or per project .omp/mcp.json), inside mcpServers:

"zenmoney": {
  "type": "stdio",
  "command": "npx",
  "args": ["-y", "mcp-server-zenmoney"],
  "env": { "ZENMONEY_TOKEN": "${ZENMONEY_TOKEN}" }
}

omp expands ${VAR} placeholders, so the token can stay in your shell environment. Then /mcp reload.

Keeping the token out of config files

umask 077 && printf '%s' 'YOUR_TOKEN' > ~/.config/zenmoney-token

and use "env": { "ZENMONEY_TOKEN_FILE": "~/.config/zenmoney-token" } instead of ZENMONEY_TOKEN.

Read-only

Add "ZENMONEY_READ_ONLY": "true" to env, or --read-only to args. Write tools are then not registered at all.

HTTP mode

ZENMONEY_TOKEN=... npx -y mcp-server-zenmoney --http --port 3000
  • MCP endpoint (Streamable HTTP): http://127.0.0.1:3000/mcp; GET /health returns {"ok":true}; any other path is 404.

  • Default bind is 127.0.0.1. Loopback binds (127.0.0.1, localhost, ::1) apply Host/Origin validation against DNS rebinding.

  • Any other --host / MCP_HTTP_HOST is refused unless MCP_HTTP_TOKEN is set. When set (on any bind), clients must send Authorization: Bearer <MCP_HTTP_TOKEN>, otherwise they get 401. It is a single-user server: whoever has that token has your ZenMoney access, so keep it on a private network.

Configuration

Variable / flag

Default

Meaning

ZENMONEY_TOKEN

—

ZenMoney API token from https://zerro.app/token

ZENMONEY_TOKEN_FILE

—

File containing the token (used when ZENMONEY_TOKEN is empty; ~ expanded)

ZENMONEY_READ_ONLY / --read-only

false

1/true/yes/on registers only read tools

ZENMONEY_SYNC_TTL

30

Seconds before data is re-synced on the next tool call

ZENMONEY_CACHE_DIR

$XDG_CACHE_HOME/mcp-server-zenmoney, else ~/.cache/mcp-server-zenmoney

Snapshot cache directory

ZENMONEY_CACHE

on

0/false/no/off keeps data in memory only

ZENMONEY_API_URL

https://api.zenmoney.ru, fallback https://api.zenmoney.app

Pin a single API origin

MCP_TRANSPORT / --http

stdio

http serves Streamable HTTP

MCP_HTTP_HOST / --host

127.0.0.1

HTTP bind address

MCP_HTTP_PORT / --port

3000

HTTP port

MCP_HTTP_TOKEN

—

Bearer token HTTP clients must send; required for non-loopback binds

--check

—

Sync once with the configured token, print a one-line summary and exit

--help, -h / --version, -v

—

Print usage / version

CLI flags take precedence over the corresponding environment variables.

Tools

Read tools (always registered, annotated readOnlyHint):

Tool

Purpose

get_overview

Start here: today's date, main currency, net worth, account balances, this month's income/expenses, upcoming planned operations, data freshness

list_accounts

Accounts with balances (also in main currency), credit limits, bank and flags

list_categories

Category tree with kind (expense/income/both) and 12-month usage counts

list_payees

Known payees ranked by use, with last date and usual category

list_transactions

Filtered, paginated transaction search plus income/expense totals of all matches

summarize_transactions

Aggregate by category, parent category, payee, account or day/week/month/year; measure=both for cash flow

get_budget

Monthly budget vs. actual, remaining, % used and still-planned operations per category

list_planned

Planned / recurring operations still pending in a date range, with expected totals

list_debts

Per-person debt balances (positive = they owe you); optionally a person's debt transactions

suggest_category

Category/merchant suggestions for payee names from ZenMoney and your own history

sync

Pull changes from ZenMoney now; full=true re-downloads everything

Write tools (not registered in read-only mode):

Tool

Purpose

create_transactions

Record up to 50 expenses, incomes, transfers or debt operations; auto-categorizes, warns about duplicates

update_transactions

Edit amount, account, category, payee, comment, date, original amount; transfer destination. Type cannot change

delete_transactions

Soft-delete transactions by id

restore_transactions

Bring deleted transactions back (as copies with new ids)

create_category

Create a category or subcategory (one nesting level)

update_category

Rename, move, change kind, budget inclusion, mandatory flag, archive

delete_category

Delete a category, moving its transactions and planned operations to move_to or uncategorizing them

set_budgets

Set exact monthly budgets per category, total or uncategorized; 0 removes

create_account

Create a cash, card, checking or e-money account

update_account

Rename, archive, include/exclude from balance, savings flag, credit limit

adjust_account_balance

Match a real balance by recording a correction income/expense

Write tools carry MCP annotations: creating tools are non-destructive (destructiveHint: false); editing/deleting tools are destructiveHint: true, so clients that honor hints can ask for confirmation.

create_transactions, update_transactions, delete_transactions, delete_category and set_budgets accept dry_run: true: they resolve everything and return the would-be result (for updates: before/after per transaction) without sending anything to ZenMoney.

Prompts

Clients that support MCP prompts (e.g. as slash commands) get three workflows that only orchestrate the tools above:

Prompt

Arguments

What it does

monthly_review

month (YYYY-MM)

Income, spending vs. previous month, budget status, large transactions, upcoming payments, suggestions

categorize_transactions

period

Proposes categories for uncategorized transactions, applies after confirmation (dry run first)

plan_budget

month (YYYY-MM)

Proposes budgets from the last three months and planned payments, applies with set_budgets after confirmation

Conventions

These are also sent to the agent as server instructions.

  • Amounts are always positive; type gives the direction: expense, income, transfer (between own accounts), debt_out (money given to a person: lending or repaying), debt_in (money received: borrowing or being repaid).

  • Accounts and categories are referenced by name (case- and emoji-insensitive, unique substrings work), "Parent / Child" category paths, or ids. Accounts can also be referenced by card last digits (4+).

  • Currencies are ISO codes (RUB, USD, KZT, ...). Dates are YYYY-MM-DD.

  • Query tools accept period: today, yesterday, this_week, last_week, this_month, last_month, this_quarter, last_quarter, this_year, last_year, last_7_days, last_30_days, last_90_days, last_12_months (weeks start Monday; date_from/date_to override either end).

  • Reports are in the main currency at ZenMoney's current exchange rates and, like ZenMoney, count only in-balance accounts unless asked otherwise. Transfers and debts are excluded from summaries. The totals returned by list_transactions follow the same rules (in-balance, not deleted), even when the listed transactions include off-balance or deleted ones.

  • A missing expense/income category on new transactions is filled from this payee's history first, then from the ZenMoney suggestion service (only live categories of the right kind). A new transaction matching an existing one (same date, type, account and amount) produces a duplicate warning; it is still recorded.

  • update_transactions category replaces the main category and keeps secondary tags; null leaves the transaction uncategorized. Moving a transaction to an account in another currency requires the new amount. Changing only amount of a same-currency transfer updates both sides unless the received amount differed (a fee), which is kept.

  • Every write syncs with ZenMoney first and runs one at a time, so edits made in the app moments earlier are not overwritten and parallel tool calls cannot clobber each other. References to categories/merchants deleted in the app are dropped from pushed objects instead of failing the whole batch.

  • If ZenMoney accepts a write but the answer is lost (timeout), the tool says the change may or may not have been saved; the next call re-syncs, so a retry gets a duplicate warning instead of silently doubling.

  • Deletes are soft. ZenMoney cannot undelete in place, so restore_transactions re-creates each transaction under a new id; restoring the same transaction again is a no-op that points at the existing copy.

  • Budgets: set_budgets sets an exact amount (lock on); 0 removes the budget. Budgets without the lock (set in the app) also include the month's planned operations, as in ZenMoney. The profile's month start day is honored. Rollover between months is not modelled.

  • Zerro's hidden 🤖 [Zerro Data] account and the reminders attached to it are ignored everywhere.

Example prompts

  • "How did October go? Biggest categories and anything unusual compared to September."

  • "Add: Magnum 12 450 KZT today from Kaspi Gold, groceries."

  • "Recategorize all Yandex Go rides from last month to Transport / Taxi."

  • "Copy this month's budgets to next month, but raise Restaurants to 80 000."

  • "Who owes me money, and how much in total?"

  • "What payments are scheduled for the next two weeks?"

  • "My cash wallet actually has 23 000; fix the balance."

Data, privacy and sync

  • The server keeps a local replica of your ZenMoney database synced via /v8/diff/: a full download on first use, incremental diffs afterwards. Data older than ZENMONEY_SYNC_TTL seconds is refreshed automatically before a tool runs; sync forces it.

  • The replica is persisted to <cache dir>/snapshot-<hash>.json (directory created 0700, file 0600). The filename is a hash of token and API origin; the token itself is never written to disk. The snapshot does contain your financial data. Set ZENMONEY_CACHE=off to keep everything in memory.

  • API origin: https://api.zenmoney.ru first; if it rejects the token before any request has succeeded, https://api.zenmoney.app is tried. ZENMONEY_API_URL pins a single origin.

  • HTTP 401 from ZenMoney means the token is no longer valid: get a fresh one at https://zerro.app/token.

  • "Today" and default dates use the process time zone; set TZ (e.g. TZ=Asia/Almaty) in the server's env if the agent host runs in a different zone.

  • Snapshots of tokens you no longer use stay in the cache directory; delete old snapshot-*.json files after rotating a token.

  • Logs go to stderr only.

Limitations

  • Planned (recurring) operations can be read but not created or edited.

  • Loan and deposit accounts cannot be created; bank-synced accounts are created by ZenMoney itself.

  • Balances are not edited directly; adjust_account_balance records a correction transaction.

  • Currency conversions in reports use current exchange rates, not historical ones.

  • Zerro's envelope budget rollover is not computed.

Development

git clone https://github.com/krax1337/mcp-server-zenmoney && cd mcp-server-zenmoney
npm install         # also builds dist/
npm run typecheck   # tsc --noEmit
npm test            # vitest; offline, against an in-repo fake ZenMoney server in test/
node dist/index.js  # run the local build

Releases: bump version in package.json and server.json, then push a vX.Y.Z tag. The Publish workflow tests, publishes to npm via Trusted Publishing (no token, with provenance) and to the MCP Registry as io.github.krax1337/mcp-server-zenmoney.

Source layout:

  • src/index.ts — entry point, stdio and HTTP transports

  • src/config.ts — environment variables and CLI flags

  • src/server.ts — MCP server and agent instructions

  • src/prompts.ts — MCP prompts (monthly_review, categorize_transactions, plan_budget)

  • src/tools/ — MCP tools (read.ts, write.ts, shared schemas and context)

  • src/zenmoney/ — API client (api.ts), sync store and cache (store.ts), ledger semantics (ledger.ts), analytics (analytics.ts), write builders (writes.ts), dates and types

License

MIT

Available Tools

22 tools
adjust_account_balanceReconcile account balanceA
Idempotent

Make an account balance match reality by recording the difference as an income or expense correction transaction (computed against freshly synced data).

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNoDefaults to today
accountYesAccount name, id or card last digits
commentNoBalance correction
categoryNoOptional category for the correction
actual_balanceYesThe real current balance in the account's currency

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish readOnlyHint=false, idempotentHint=true, and destructiveHint=false, matching the 'make it match reality' framing. The description adds real value beyond them: it discloses that a correction transaction is created and that computation depends on freshly synced data. It doesn't cover permissions or what happens for an unknown account, keeping it out of 5 territory.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the core action leads and the qualifying clause follows. Every part earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an idempotent mutation tool with no output schema, the description conveys the mechanism, the sync dependency, and the record type created. It leaves minor gaps around permissions and side effects, but nothing essential to invoking it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 80%, with account, actual_balance, date, comment, and category all documented inline. The description adds only the notion that the correction is income-or-expense in nature, which loosely relates to category but gives no syntax or format detail beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (adjust/reconcile the account balance) and explains the mechanism: recording the difference as an income/expense correction transaction. This is more than a restatement of the name, though it doesn't explicitly contrast itself with create_transactions, which could also produce a correcting entry.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear condition for use ('make an account balance match reality') and a prerequisite hint ('computed against freshly synced data'), which nudges the agent toward the sync tool first. It stops short of naming alternatives or stating when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_accountCreate accountB

Create a cash, card, checking or e-money account. Bank-synced accounts are created by ZenMoney itself.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNocash, ccard (card), checking, emoneycash
titleYes
balanceNoOpening balance
savingsNo
currencyNoISO code; defaults to the main currency
in_balanceNoCount in the total balance and reports
credit_limitNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, and idempotentHint=false, so the agent knows this is a non-destructive write that may duplicate on repeat calls. The description adds only the bank-sync exclusion and says nothing about permissions or currency defaults beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the supported types front-loaded and the exclusionary note second. Nothing wastes space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation tool with no output schema and only 57% schema coverage, the description is thin. It never explains opening balance behavior, savings, credit_limit, or in_balance, so an agent must infer semantics from terse schema strings alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 57%, with type, balance, currency, and in_balance already described in the schema. The description usefully glosses 'ccard' as 'card' and lists the valid types, but leaves title, savings, and credit_limit (including that credit_limit applies to card/e-money accounts) unaddressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (create) and resource (account) plus the supported account types, which maps directly to the enum. It also draws a scope boundary by excluding bank-synced accounts, though it does not name sibling tools like update_account or adjust_account_balance.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The note that bank-synced accounts are created by ZenMoney itself is a genuine when-not-to-use signal. However, there is no guidance on choosing between create_account, update_account, and adjust_account_balance, or on prerequisites such as required currency context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_categoryCreate categoryA

Create a category, optionally as a subcategory of a top-level category (ZenMoney allows one nesting level).

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoDefaults to the parent's kind, else expense
titleYes
parentNoTop-level parent category (name or id)
requiredNoMark as a mandatory (non-discretionary) expense

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare write (readOnlyHint=false), non-destructive, non-idempotent, closed-world behavior, so the safety profile is covered. The description adds a genuine domain constraint that the annotations cannot express: ZenMoney permits only one nesting level. It still says nothing about duplicate handling, permission requirements, or the creation response.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the nesting constraint is stated inline where it is most useful. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description need not explain returns, and annotations cover the mutation profile. Still, for a write tool with a parent-reference parameter, it omits how an unknown parent is handled, whether titles must be unique, and what the caller receives back – gaps an agent would hit in practice.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is high (75%) and the schema already documents parent as 'Top-level parent category (name or id)' and kind as defaulting to the parent's kind. The description reinforces the one-level nesting rule and the optionality of parent, but adds little beyond the schema; baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a category') and adds a scoping constraint ('optionally as a subcategory of a top-level category'). This clearly separates it from create_account, create_transactions, and update_category, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the verb 'Create' and the option to nest, which lets an agent infer this is for new categories rather than update_category. However, there is no explicit when-to-use/when-not guidance, no prerequisites, and no mention of what happens if the category already exists.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_transactionsAdd transactionsA

Record one or more transactions (expense, income, transfer between own accounts, debt_out / debt_in with a person). Amounts are positive in the account currency. Without a category, one is picked from this payee's history or ZenMoney suggestions (auto_categorize). Returns the created transactions, updated balances and possible-duplicate warnings.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNoPreview the outcome without saving anything
transactionsYes
auto_categorizeNoFill a missing expense/income category from history or ZenMoney suggestions

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare the mutation profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), so the description's added value is the return contract — created transactions, updated balances, duplicate warnings — and the fact that categories are auto-filled from payee history or ZenMoney suggestions. That is meaningful disclosure beyond the annotations, though batch size limits are not mentioned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with what is recorded, followed by amount/category rules and then the return contract. No filler and nothing buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly describes return values (created transactions, balances, duplicate warnings) and covers categorization behavior. Gaps are minor: the 50-item batch cap and dry_run preview are left to the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

At 67% schema coverage the description still adds real meaning: amounts are always positive in the account currency, transfers are between the user's own accounts, and category resolution falls back to payee history/suggestions when omitted. It complements rather than repeats the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (Record) plus resource (transactions) with the full set of supported types enumerated, including a plain-language gloss of debt_out vs debt_in. An agent can immediately tell this creates ledger entries rather than accounts or budgets.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied (this is the tool that writes transactions) and the auto-categorization default is stated, but no sibling is named as the alternative for editing or reversing entries (e.g. update_transactions vs delete_transactions). No explicit when-to-use or prerequisite guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_categoryDelete categoryA
DestructiveIdempotent

Delete a category. Its transactions and planned operations are moved to move_to (merge categories) or, with uncategorize=true, left without this category. Subcategories must be moved or deleted first.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNoPreview the outcome without saving anything
move_toNoCategory that takes over its transactions
categoryYesCategory name, path or id
uncategorizeNoConfirm removing the category from its transactions and planned operations when move_to is not given

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is covered. The description adds genuinely useful context beyond that: where transactions and planned operations go (moved to move_to or left uncategorized), and the subcategory prerequisite. It stops short of stating irreversibility or permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and then the consequential behavior, with no redundant restatement of the tool name or title. Every sentence conveys a distinct constraint.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation with annotations covering the safety profile and no output schema, the description covers the important effects and prerequisites. Minor gaps remain around dry_run's preview semantics and whether the deletion is reversible, but the essential call-time information is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real semantics beyond the schema: it explains the interplay between move_to and uncategorize=true (uncategorize only applies when move_to is absent) and frames move_to as the merge mechanism. That is meaning the schema text does not supply.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Delete a category') that an agent can immediately distinguish from siblings create_category and update_category. The following sentences immediately qualify the scope of the deletion, reinforcing what the tool actually does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains the conditions that govern deletion behavior: choose move_to to merge, or use uncategorize=true when no move_to is given, and it states the prerequisite that subcategories must be moved or deleted first. It does not name a sibling tool as an alternative route, but the when-to-use conditions are clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_transactionsDelete transactionsA
DestructiveIdempotent

Delete transactions by id (soft delete; restore_transactions brings them back).

ParametersJSON Schema
NameRequiredDescriptionDefault
idsYes
dry_runNoPreview the outcome without saving anything

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=true and idempotentHint=true; the description adds the genuinely useful nuance that this is a soft delete reversible via restore_transactions, which prevents an agent from treating destruction as permanent. It does not mention the 100-id batch ceiling, partial-failure behavior, or that dry_run is available.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with the verb and resource front-loaded, followed immediately by the reversibility caveat and its sibling pointer. No filler, nothing buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter mutation with no output schema, the description plus annotations cover the essentials: what it does, that it is destructive, idempotent, and reversible. The main gap is the batch limit and preview option, which the agent can still discover from the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%: dry_run is documented in the schema, but ids is not. The description only confirms deletion is keyed by id, adding no detail on the array form or the 100-item limit, and says nothing about dry_run. It partially compensates but leaves the undocumented parameter unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Delete transactions') plus the keying mechanism ('by id'). The parenthetical distinguishes the operation from its sibling restore_transactions, so an agent can route between them without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description names restore_transactions as the undo path, giving clear context for the reversibility of this action. It stops short of saying when not to use this tool (e.g. use update_transactions to modify instead) or any preconditions, so it is clear but not fully exclusive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_budgetMonthly budget statusB
Read-only

ZenMoney budget for a month: per category budget, actual, remaining, % used and still-planned operations, with subcategories and totals, in the main currency. Uses the profile's month start day. Rollover between months is not modelled.

ParametersJSON Schema
NameRequiredDescriptionDefault
sideNoexpense
monthNoYYYY-MM; defaults to the current month

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds genuinely useful behavior beyond that: values are in the main currency, the month start day is taken from the profile, and rollover between months is explicitly not modelled. These are non-obvious traits that materially affect interpretation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the resource and then lists the payload fields, followed by two short constraint clauses. No filler, though the enumeration is tightly packed and could be marginally clearer.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining returns and does so (per-category, subcategories, totals, currency). Read-only safety is covered by annotations, and the rollover caveat closes the main interpretive gap; only the side parameter semantics remain unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: month is documented in the schema but side is an unexplained enum (expense/income/both). The description does not clarify what side does or how it changes the result, so it fails to compensate for the coverage gap on the undocumented parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb+resource ('ZenMoney budget for a month') and enumerates the returned fields (budget, actual, remaining, % used, planned operations, subtotals), so the agent knows exactly what this tool produces. It does not, however, explicitly distinguish itself from its nearest siblings like set_budgets (the write counterpart) or get_overview, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance and no reference to any alternative tool. The description only elaborates what is returned and a couple of behavioral quirks; it never states the condition that should make an agent pick get_budget over get_overview or summarize_transactions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_overviewFinancial overviewA
Read-only

Start here. Returns today's date, the user's main currency, net worth, active account balances, this month's income/expenses with top categories, upcoming planned operations and data freshness.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds real value beyond that by enumerating the composite payload (net worth, balances, categories, planned operations) and flagging "data freshness" as a returned signal, though it doesn't discuss caching or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero filler; the routing cue ("Start here") is front-loaded before the payload enumeration. Every clause earns its place and maps to actual returned content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters, no output schema, and annotations covering the safety profile, the description's full enumeration of the return payload is sufficient for an agent to decide to call it and know what to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the rubric the baseline is 4. There are no inputs whose semantics need explaining.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific read action and enumerates exactly what the resource returns: date, currency, net worth, balances, income/expenses with categories, planned operations, data freshness. This clearly differentiates it from siblings like list_accounts and summarize_transactions, which return narrower slices of the same domain.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Start here" is an explicit usage directive that positions this as the first call an agent should make, which is unusually strong guidance. It stops short of naming alternatives or stating when-not to use it, so it falls just below a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_accountsList accountsA
Read-only

Accounts with balances, currency, credit limit, bank and flags. balance_main is the balance converted to the main currency. Totals cover in-balance accounts. Debt bookkeeping is in list_debts.

ParametersJSON Schema
NameRequiredDescriptionDefault
typesNocash, ccard (card), checking, loan, deposit, emoney
searchNoSubstring of the account name or card number
include_archivedNoAlso return archived accounts

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, and the description goes further by disclosing domain behavior: balance_main is the main-currency conversion and totals cover only in-balance accounts. It does not cover return ordering or pagination, but the semantic disclosure is genuinely additive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four terse sentences, each carrying distinct information (fields, conversion semantics, totals scope, sibling routing). It opens with a field list rather than a purpose statement, which is slightly less front-loaded than ideal, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description rightly explains the returned fields, the meaning of balance_main, and the scope of totals. Combined with full schema coverage, an agent has enough to call it correctly; only filtering/archived behavior nuances are left implicit.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so types, search, and include_archived are already documented in the schema. The description adds no parameter-level meaning (e.g., how search interacts with archived filtering), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description makes clear this returns account records and enumerates the fields carried (balances, currency, credit limit, bank, flags), and it explicitly routes debt data elsewhere. The verb is only implied by the tool name/title rather than stated, but the resource and scope are unambiguous and distinguished from sibling list_debts.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names one exclusion ('Debt bookkeeping is in list_debts'), which gives partial routing, but offers no when-to-use guidance relative to siblings like get_overview, create_account, or update_account. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_categoriesList categoriesA
Read-only

Category tree (ZenMoney "tags"; one nesting level) with ids, kind (expense/income/both) and number of uses in the last 12 months. Use names, "Parent / Child" paths or ids elsewhere.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoOnly categories usable for this kind of transactionall
searchNoSubstring of the category name
include_archivedNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered; the description adds real value by disclosing the tree shape (one nesting level), the 'tags' alias, and that usage counts cover the last 12 months.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with the resource and return shape front-loaded and no filler; the trailing usage note earns its place by linking to sibling tools.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description does the necessary work of describing the returned tree, ids, kinds and 12-month usage counts, but it leaves include_archived and the exact kind enum behavior unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%, and the description maps kind to expense/income/both while the schema enum is expense/income/all — a minor mismatch — and never explains include_archived or the search substring filter. It adds path/name referencing semantics but not full parameter clarity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the specific resource (category tree / ZenMoney 'tags'), its one-level nesting, and the fields returned (ids, kind, usage count over 12 months). It clearly differs from create/update/delete_category siblings, though it doesn't explicitly say so.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The sentence 'Use names, "Parent / Child" paths or ids elsewhere' tells the agent how to reference categories in other tools, which implicitly frames this as the lookup step, but there is no explicit when-to-use/when-not guidance or named alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_debtsDebts by personA
Read-only

Who owes whom: per-person debt balances from debt operations (positive = they owe you, negative = you owe them), per currency and in the main currency. Pass person to also get their debt transactions.

ParametersJSON Schema
NameRequiredDescriptionDefault
personNoSubstring of the person name
include_settledNoInclude people whose debts net to zero

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description adds meaningful output semantics (positive = they owe you, negative = you owe them, per currency and main currency) that annotations cannot carry. It does not disclose pagination, sorting, or whether balances are computed live.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler; the core purpose and sign convention are front-loaded, and the parameter behavior is appended last where it belongs. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the burden of explaining return values, and it does so well via the sign convention and currency dimension. It is only mildly incomplete: it says nothing about ordering, empty results, or how settled people appear by default.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning the schema lacks: passing 'person' also returns that person's debt transactions, which is not implied by the schema's 'Substring of the person name'. The interaction with include_settled is left to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource and semantics: per-person debt balances derived from debt operations, with the sign convention spelled out. It clearly distinguishes itself from transaction/account/category siblings by domain. It stops short of naming an alternative sibling, but the resource is unique enough that confusion is unlikely.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the resource ('who owes whom'), and the description explains what passing 'person' adds, which is a light form of usage guidance. However, there is no explicit when-to-use/when-not, no mention of an alternative for raw transaction history, and no note on when to set include_settled.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_payeesList payeesA
Read-only

Known payees/merchants ranked by number of transactions, with last date and the category usually used for them. Useful for consistent categorization and payee spelling.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
searchNoSubstring of the payee name

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds genuine behavioral context (results are ranked by transaction frequency and include last-used date and usual category), but says nothing about ordering guarantees, pagination, or how the default limit truncates results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences; the payload description (what fields come back) is front-loaded and the benefit statement follows. Nothing is redundant, though the closing sentence is soft and could have carried the limit/ordering semantics instead.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no output schema, the description usefully enumerates the returned fields (payee, rank, last date, usual category). Missing only parameter-level behavior (default cap of 50/500 max, sort direction), which keeps it just short of complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: 'search' is documented in the schema but 'limit' is not. The description does not mention either parameter, so it adds nothing beyond the schema. Because limit carries a visible default of 50 and min/max bounds it is largely self-explanatory, keeping this at a baseline 3 rather than lower.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (lists known payees/merchants) and adds differentiating detail: ranked by transaction count, with last date and typical category. It is clearly distinct from list_accounts, list_categories, and suggest_category, though it never names a sibling explicitly to contrast with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Useful for consistent categorization and payee spelling' implies the use case but gives no explicit when-to-use/when-not guidance and does not point to an alternative such as suggest_category. Usage must be inferred from context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_plannedPlanned operationsA
Read-only

Scheduled / recurring operations (ZenMoney reminders) that are still planned in a date range, with recurrence, overdue and forecast flags, plus expected expense/income totals.

ParametersJSON Schema
NameRequiredDescriptionDefault
date_toNoDefaults to 30 days ahead
date_fromNoDefaults to 7 days ago (to surface overdue items)
include_forecastNoInclude ZenMoney's auto-generated forecast entries

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so the safety profile is covered. The description usefully adds that results include recurrence, overdue and forecast flags plus expected expense/income totals, but says nothing about pagination, result size, or how the overdue/forecast flags are computed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the resource and then the returned signals; nothing is wasted, though the list-of-nouns style is slightly heavy for one sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only list tool with no output schema, the description does carry the return-shape burden by naming the flags and totals. Only minor gaps (pagination/volume, ordering) remain.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the defaults (30 days ahead, 7 days ago, forecast included) are documented per-parameter in the schema. The description adds only the date-range concept, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (scheduled/recurring ZenMoney reminders still planned in a date range), and enumerates the payload content (recurrence, overdue, forecast flags, expected totals). It is clearly distinguishable from siblings like list_transactions or list_debts, though it never names a contrasting sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Use is only implied by the phrase 'still planned in a date range'; there is no explicit statement of when to prefer this over list_transactions or get_budget, and no prerequisite or exclusion guidance. Adequate but leaves the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_transactionsSearch transactionsA
Read-only

Find transactions with filters (dates/period, type, accounts, categories incl. subcategories, payee, free text, amount, currency). Newest first by default, paginated. Also returns income/expense totals of all matches in the main currency (in-balance accounts only, like ZenMoney reports), so "how much did I spend on X" needs only this call (or summarize_transactions for breakdowns).

ParametersJSON Schema
NameRequiredDescriptionDefault
idsNoFetch specific transactions by id
sortNodate_desc
limitNo
payeeNoSubstring of the payee / merchant name
typesNoTransaction types to include: expense | income | transfer (between own accounts) | debt_out (money given to a person: lending, or repaying them) | debt_in (money received from a person: borrowing, or them repaying you)
offsetNo
periodNoDate range relative to today (weeks start on Monday). date_from/date_to override either end.
searchNoFree text matched against payee, merchant, comment and category names
date_toNoInclusive end date, YYYY-MM-DD
accountsNoAccount names, ids or card last digits; matches money leaving or arriving
currencyNoOnly transactions in this currency (ISO code)
date_fromNoInclusive start date, YYYY-MM-DD
categoriesNoCategory names, "Parent / Child" paths or ids; subcategories are included
max_amountNoMaximum amount in the account's currency
min_amountNoMinimum amount in the account's currency
uncategorizedNoOnly transactions without a category
include_deletedNoInclude deleted transactions (they can be restored)

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and closed-world behavior, so the safety profile is covered. The description adds genuinely useful behavior beyond annotations: default sort order (newest first), pagination, and that income/expense totals are computed over ALL matches in the main currency and only for in-balance accounts. It does not describe the transaction object shape or page-size limits, keeping this at a 4.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with the core verb and filter list front-loaded, followed by output-behavior context. Every clause carries information; the parenthetical about summarize_transactions is slightly compressed but earns its place as routing guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter, no-output-schema tool with minimal annotations, the description covers the essential gaps: default ordering, pagination, and the aggregation semantics of the returned totals. What is missing is the shape/content of the returned transactions themselves, but the totals explanation handles the highest-risk ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 82%, so the schema already documents nearly all 17 parameters, including enum definitions and formats. The description's parameter recap (dates/period, type, accounts, categories, payee, free text, amount, currency) largely restates what the schema already provides, so it earns the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Find transactions') and enumerates the filter dimensions, so the agent knows exactly what it retrieves. It also explicitly distinguishes itself from summarize_transactions and clarifies that it returns totals itself, which differentiates it from siblings without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete routing guidance: '"how much did I spend on X" needs only this call (or summarize_transactions for breakdowns)', directing the agent to the right alternative for breakdowns. It stops short of stating when this tool should NOT be used, but the usage context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

restore_transactionsRestore deleted transactionsA
Idempotent

Bring back deleted transactions (find them with list_transactions include_deleted=true). ZenMoney cannot undelete in place, so each comes back as a copy with a new id; already-restored ones are skipped.

ParametersJSON Schema
NameRequiredDescriptionDefault
idsYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds substantial behavior beyond annotations: ZenMoney cannot undelete in place, restored items come back as copies with new ids, and already-restored items are skipped. This is consistent with idempotentHint=true and destructiveHint=false, and discloses the surprising id-changing side effect an agent would otherwise not anticipate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence front-loads the action, then appends the lookup hint and the key behavioral caveat. No filler; every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the mutation, its idempotency, and the new-id caveat, which is the critical surprise. It does not describe the return payload or partial-failure behavior, but with no output schema and a single-parameter surface, the description is nearly sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter with 0% schema description coverage, so the description carries some burden. It implies the transactions to restore and how to obtain their ids, but never states that 'ids' is an array of up to 100 transaction identifiers, leaving the shape to the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Bring back deleted transactions') and immediately distinguishes itself from the sibling delete_transactions and list_transactions. An agent can tell exactly what this does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the precondition and the alternative for locating inputs ('find them with list_transactions include_deleted=true'), which is actionable routing guidance. It does not state explicit when-not-to-use cases, but the usage context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_budgetsSet monthly budgetsA
DestructiveIdempotent

Set exact monthly budget amounts (main currency) per category, for "total" (overall monthly budget) or "uncategorized". 0 removes that budget. Batch-friendly: e.g. copy last month's budgets to next month.

ParametersJSON Schema
NameRequiredDescriptionDefault
budgetsYes
dry_runNoPreview the outcome without saving anything

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, covering the safety profile. The description reinforces this by explaining that a 0 value removes the budget, which is meaningful destructive semantics, but it omits permission requirements and the effect of partial failures in a batch.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core semantics ('set exact monthly budget amounts per category') front-loaded. The batch example is useful but slightly loose; nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, idempotent mutation with no output schema and no annotations covering permissions, the description covers the value semantics but omits auth/permission requirements and any confirmation-of-outcome behavior beyond dry_run (which is only in the schema). Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50% and the description adds real meaning not in the schema: amounts are 'exact', denominated in the main currency, and applied per category. It does not mention dry_run, which the schema already documents, so it is not fully compensating.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (set) and resource (monthly budgets per category), and specifies the special category targets 'total' and 'uncategorized'. An agent can distinguish this from the sibling get_budget without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a concrete usage scenario ('copy last month's budgets to next month') and the batch-friendly framing, so the agent knows the intended bulk workflow. It does not name an alternative tool or state when not to use this one, so it falls short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

suggest_categorySuggest category for payeesA
Read-only

Category and merchant suggestions for payee names, from ZenMoney's suggestion service and from this user's own history.

ParametersJSON Schema
NameRequiredDescriptionDefault
payeesYesPayee names as they appear on a receipt or statement

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds genuine context about provenance (ZenMoney's suggestion service plus the user's own history), which explains why results may vary per user. It does not mention cost, rate limits, or what happens when no suggestion is found.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the two suggestion sources are packed into one clause. It is appropriately sized for a simple one-parameter read tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only suggestion tool with no output schema and a fully documented single parameter, the description gives enough to call it correctly and understand where answers come from. It stops short of describing result shape or fallback behavior, which is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there is a single parameter, so the schema already explains that payees are names as they appear on a receipt or statement. The description adds no format, casing, or matching-behavior guidance beyond that, so it sits at the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (suggest) and resource (category and merchant for payee names), so the agent knows exactly what the tool produces. It does not explicitly contrast itself with siblings like list_categories or list_payees, but the 'suggestion' framing is distinctive enough to separate it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: an agent can infer it should call this when it has payee names and needs a category/merchant guess. There is no statement of when to prefer this over list_categories or a manual lookup, and no prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

summarize_transactionsSpending / income reportA
Read-only

Aggregate expenses and/or income by category, parent category, payee, account or time (day/week/month/year) with totals, counts and shares. measure=both gives income, expense and net per group (cash flow). Transfers and debts are excluded; amounts are converted with current rates. By default only in-balance accounts count, like ZenMoney reports.

ParametersJSON Schema
NameRequiredDescriptionDefault
topNoMax groups for non-time groupings; the rest is rolled up
payeeNoSubstring of the payee / merchant name
typesNoTransaction types to include: expense | income | transfer (between own accounts) | debt_out (money given to a person: lending, or repaying them) | debt_in (money received from a person: borrowing, or them repaying you)
periodNoDate range relative to today (weeks start on Monday). date_from/date_to override either end.
searchNoFree text matched against payee, merchant, comment and category names
date_toNoInclusive end date, YYYY-MM-DD
measureNoexpense
accountsNoAccount names, ids or card last digits; matches money leaving or arriving
currencyNoOnly transactions in this currency (ISO code)
group_byNocategory
date_fromNoInclusive start date, YYYY-MM-DD
categoriesNoCategory names, "Parent / Child" paths or ids; subcategories are included
max_amountNoMaximum amount in the account's currency
min_amountNoMinimum amount in the account's currency
uncategorizedNoOnly transactions without a category
report_currencyNoISO code to report in; defaults to the main currency
include_off_balanceNoAlso count accounts excluded from the balance

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true and openWorldHint=false, so the description carries the behavioral load and does it well: exclusion of transfers/debt types, currency conversion at current rates, and the default in-balance account filter with an override. It stops short of describing the response shape (groups, totals, counts, shares ordering) or the effect of the top roll-up.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core capability and grouping options, with no filler. It is dense and each clause carries information, though the trailing 'like ZenMoney reports' reference adds minimal operational value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 17-parameter tool with no output schema, the description covers the key semantics an agent needs: what gets aggregated, what is excluded, currency conversion, default account scoping, and that the output contains totals, counts and shares. Remaining gaps (how top groups are rolled up, exact return structure) are modest.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is high (88%), so the baseline is 3, but the description adds meaning the schema lacks — notably the semantics of measure=both (income + expense + net cash flow) and the relationship between include_off_balance and the default in-balance filtering. The undocumented enum for measure is effectively compensated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (aggregate) and resource (expenses/income) plus the full set of grouping dimensions (category, parent category, payee, account, time), which clearly separates it from the raw-listing siblings like list_transactions and from get_overview/get_budget. An agent can identify it as the reporting/aggregation tool without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides real selection context: measure=both yields income, expense and net per group (cash flow), transfers and debts are excluded, and only in-balance accounts count by default unless include_off_balance is set. It never explicitly names an alternative tool (e.g. list_transactions for raw rows), so it is clear context without exclusion-vs-sibling routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

syncSync with ZenMoneyA
Read-onlyIdempotent

Pull the latest changes from ZenMoney now (tools already auto-sync when data is older than the configured TTL). full=true re-downloads everything.

ParametersJSON Schema
NameRequiredDescriptionDefault
fullNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and closed-world. The description adds the crucial non-obvious fact that sync is automatic, so a caller knows this is a manual override – behavior not captured by any annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight clauses: one naming the action and the auto-sync caveat, one defining the boolean. Every sentence earns its place and the primary purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single optional boolean with no output schema, the description gives what's needed: what it does, that it's normally automatic, and what full does. Only minor gap is that it doesn't say what 'pull' returns or whether concurrent calls are safe, but annotations cover safety.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the only param is 'full' with no description. The description supplies the semantics: full=true re-downloads everything, versus an incremental pull by default. This directly compensates for the undocumented param.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (pull) and resource (changes from ZenMoney), with a clear action. Doesn't explicitly compare to siblings, but the 'auto-sync' note distinguishes it from the read tools in the list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The parenthetical explains that syncing already happens automatically after TTL expiry, which tells the agent when this tool is actually needed (to force a refresh). No explicit exclusions, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_accountEdit accountA
DestructiveIdempotent

Rename, archive/unarchive, include in or exclude from balance, mark as savings, or change the credit limit. To fix a balance use adjust_account_balance.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleNo
accountYesAccount name, id or card last digits
savingsNo
archivedNo
in_balanceNo
credit_limitNo

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=true, idempotentHint=true, and openWorldHint=false, so the safety profile is known without the description. The description's field list implies the mutation surface but adds no behavioral context beyond it — it never says whether omitted fields are left unchanged, whether archiving removes the account from reports, or whether the call is immediately persisted.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences: the operation list is front-loaded and the disambiguation follows immediately. No filler, no restatement of the name or title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, idempotent mutation tool with no output schema and six parameters, the description covers the operation set and routes the balance-correction case correctly. The remaining gap is that it doesn't state whether unspecified fields remain untouched or what permissions are required, but annotations and the operation list make it callable as-is.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 17% (just 'account'), so the description has to carry the parameters, and it largely does: rename→title, archive/unarchive→archived, include/exclude in balance→in_balance, mark as savings→savings, change credit limit→credit_limit. It stops short of saying these are the only optional fields or that they can be combined in a single call, but the mapping is a genuine gain over bare schema names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource and enumerates the concrete mutations it performs (rename, archive/unarchive, include/exclude from balance, mark as savings, change credit limit), so an agent knows exactly which field-level edits are in scope. It also explicitly contrasts itself with adjust_account_balance, which handles balance correction — a sibling it would otherwise be confused with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The closing sentence gives an explicit alternative with its triggering condition: use adjust_account_balance to fix a balance rather than this tool. That is real when-to-use guidance, though there is no broader statement of exclusions or prerequisites (e.g. required permissions, whether multiple fields can be combined in one call).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_categoryEdit categoryB
DestructiveIdempotent

Rename, move (parent null = top level), change kind, budget inclusion, mandatory flag or archive a category.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
titleNo
parentNoNew top-level parent; null makes it top-level
archivedNo
categoryYesCategory name, path or id
requiredNo
include_in_budgetNoWhether the category counts in budget calculations

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare write/destructive/idempotent behavior, so the safety profile is covered. The description adds the useful 'parent null = top level' convention, but omits that updates are partial (unspecified fields are left unchanged) and does not clarify what makes the operation destructive (e.g., archiving).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence that front-loads the operations and inlines the null-parent convention. Every clause maps to a real parameter, with only mild readability cost from the parenthetical nesting.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter mutation tool with no output schema, the description covers which fields can change but not the return behavior, partial-update semantics, or the practical effect of archiving. It is adequate but leaves real gaps an agent would want closed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema coverage at 43%, the description compensates by mapping friendly names to ambiguous fields: 'rename' → title, 'mandatory flag' → required, 'budget inclusion' → include_in_budget, 'archive' → archived. This resolves the two undocumented-looking flags, though it adds nothing beyond the schema for kind's enum values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The verb 'edit/rename/move/change/archive' is paired with the resource 'category', and the description enumerates every mutable aspect (title, parent, kind, budget inclusion, required flag, archived). An agent can tell it apart from create_category and delete_category without opening a schema, though it never names those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given: it does not say to prefer this over delete_category for removals, nor that it only updates the fields supplied. Usage is left entirely to inference from the enumerated operations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_transactionsEdit transactionsA
DestructiveIdempotent

Change one or more existing transactions: amount, account, main category (null = uncategorized), payee, comment, date, original currency amount; transfers also to_account / to_amount. The type cannot change (delete and re-create instead). Use list_transactions to find ids.

ParametersJSON Schema
NameRequiredDescriptionDefault
dry_runNoPreview the outcome without saving anything
updatesYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered. The description adds the constraint that type cannot change, but it doesn't expand on the destructive nature or preview behavior beyond what dry_run's schema description says. It does not contradict the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three compact sentences front-load the operation, list editable fields, state the immutability constraint, and route to list_transactions. No filler or redundant restatement of the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-param, single-object mutation with nested item fields, the description covers fields, null behavior, transfer nuances, and the immutability rule. It lacks specifics on the destructive/idempotent behavior implications, but annotations and dry_run's schema description fill those gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%, so the description compensates by naming transfer-specific fields (to_account, to_amount) and explaining null semantics for category ('null = uncategorized'). It doesn't restate id/date/amount syntax, but the field-level semantics it supplies exceed the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Change one or more existing transactions') and immediately enumerates editable fields, distinguishing it from create_transactions, delete_transactions, and restore_transactions. It also names exactly what cannot change (type) and points to delete/re-create as the alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly directs the agent to use list_transactions to find ids, which is the key prerequisite for invocation, and it tells the agent to delete and re-create when the type must change. It doesn't state when to prefer this tool over create_transactions for bulk edits, but the guidance is materially useful.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 22 tool updatesv1.0.0
    • First observedadjust_account_balance
    • First observedcreate_account
    • First observedcreate_category
    • First observedcreate_transactions
    • First observeddelete_category
    • First observeddelete_transactions
    • First observedget_budget
    • First observedget_overview
    • First observedlist_accounts
    • First observedlist_categories
    • First observedlist_debts
    • First observedlist_payees
    • First observedlist_planned
    • First observedlist_transactions
    • First observedrestore_transactions
    • First observedset_budgets
    • First observedsuggest_category
    • First observedsummarize_transactions
    • First observedsync
    • First observedupdate_account
    • First observedupdate_category
    • First observedupdate_transactions

TDQS

A3.8/5.0

Scored across 22 tools

Disambiguation4/5

Most tools target clearly distinct resources and actions, and descriptions explicitly disambiguate overlaps (e.g., list_transactions vs. summarize_transactions). However, several tools provide overlapping aggregate views (get_overview, list_transactions totals, get_budget, summarize_transactions) that an agent could confuse for a given reporting question.

Naming Consistency4/5

The set follows a consistent snake_case verb_noun pattern (list_accounts, create_transactions, update_category, delete_transactions, etc.). The lone bare verb 'sync' is a minor deviation, but the convention is otherwise predictable and readable.

Tool Count4/5

22 tools is on the heavy side for the 3-15 sweet spot, but each tool maps to a distinct finance operation (accounts, categories, transactions, budgets, planned, debts, analytics) and the domain genuinely supports this breadth. It is slightly over-scoped rather than bloated.

Completeness4/5

Coverage is strong: full CRUD for transactions (including restore), categories, and accounts, plus budgeting, debt, and analytics views. Gaps remain for planned/recurring operation creation and editing (only list_planned exists) and payee management (only list/suggest), which agents must work around.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to interact with YNAB budgets through natural language. Supports managing accounts, categories, transactions, and budget months with 21 tools for comprehensive budget operations.
    -
  • A
    license
    A
    quality
    D
    maintenance
    Enables AI assistants to query and manage personal finance data through the unofficial Money Lover REST API. It provides 27 tools covering authentication, wallets, categories, transactions, events, debts, and static configuration with both read and write capabilities.
    33
    17 npm
    ISC
  • A
    license
    B
    quality
    C
    maintenance
    Enables AI agents to interact with your Lunch Money personal finance data, providing tools for managing transactions, categories, budgets, assets, and accounts.
    15
    35 npm
    ISC
  • F
    license
    C
    quality
    B
    maintenance
    Enables AI agents to access and analyze financial data from Toshl Finance, including accounts, categories, budgets, and entries, through MCP resources and tools.
    27
    4
    -