Skip to main content
Glama
udaysrinu

ExpensifyAI

by udaysrinu

ExpensifyAI

Talk to your Splitwise. Get CRED-grade spending analytics.

ExpensifyAI is a Model Context Protocol (MCP) server for Splitwise that lets any LLM client (Claude, etc.) manage shared expenses and produce deterministic, premium spending analytics — category breakdowns, monthly trends, per-member comparisons, and minimum-transaction settlement plans — rendered as a self-contained, offline HTML dashboard.

Built on top of the excellent tarunn2799/splitwise-mcp; extended with a deterministic analytics engine and a category-first dashboard.

ExpensifyAI dashboard

Interactive dashboard, generated from synthetic data. Open examples/demo-dashboard.html in a browser to try it live — pick a date range and every section recomputes instantly.

Analytics (what makes this ExpensifyAI)

  • Deterministic by construction — every number is computed in pure Python with Decimal math (no float drift, no LLM estimation). Same input → byte-identical output. Each report carries a reconciliation check: per-expense shares must sum to cost, or the mismatch is flagged.

  • Category-first, à la CRED — an expandable "where it goes" view leads every report; tap a category to drill into its transactions.

  • Seven analytics modules — category breakdown · monthly trend · owed-vs-paid ("mine vs split") · per-member comparison + category×member matrix · transaction ledger · top transactions · settlement optimizer (minimum transactions to settle a group — no other Splitwise tool has this).

  • Interactive dashboard — CRED-grade dark UI with a live date-range picker + presets (this month / 3mo / 6mo / this year / all) that re-filter and recompute every section in the browser. Hand-rolled inline-SVG charts, validated colorblind-safe palette, fully offline (self-contained single file — no CDN, no server). All client math is integer paise, so the live recompute stays exact and reconciles against the Python source of truth.

  • Two analytics tools — analyze_spending(target_type, target_id?, dates?, generate_dashboard?) and compare_group_members(group_id, …). target_type is me | group | friend.

Related MCP server: Splitwise MCP Server

Splitwise's own search is Pro-paywalled and the API has no search endpoint. ExpensifyAI mirrors your data into a local SQLite DB (~/.expensifyai/splitwise.db) and searches it offline.

  • sync_all(full=False) — delta sync: uses the API's updated_after cursor so only expenses added/edited/moved/deleted since the last sync are fetched (first run pulls everything; a re-sync with no changes is a single call). Upserts by expense id, so it's idempotent. Groups and friends are fully refreshed each run (small). Deletes are mirrored (soft-deleted, excluded from search by default).

  • search_expenses(query?, min_amount?, max_amount?, user_id?, group_id?, category?, dates?) — full-text (FTS5) over description/details/category plus structured filters, against the local DB. Instant, offline, covers your entire history across all groups. Beats the paywalled app search.

Each expense's full raw API object is stored (raw column) as the source of truth, so displayed values stay faithful — the indexed REAL columns are only for filtering.

Itemization, receipt scanning & default splits

Splitwise-Pro-parity, built deterministically:

  • Structured itemization — create_itemized_expense(description, group_id, items, …) turns line-items into ONE expense where each item can split differently (beers ¾ to one person, groceries 4-way, cake between two). Each person's total owed_share is computed in exact integer paise (largest-remainder rounding, so an indivisible ₹100/3 still sums back to ₹100), and the expense is reconciled to its total before anything is written — a mismatch refuses to create rather than posting a wrong split. dry_run=True previews the computed split.

  • Receipt scanning (LLM-vision-native) — no OCR engine, no cloud keys, no new dependencies: the calling agent (Claude) reads the receipt image, extracts line-items, and calls create_itemized_expense. The server owns the exact math and the Splitwise write.

  • Receipt image upload — attach_receipt(expense_id, image_path) uploads a local image/PDF to an existing expense (multipart), so the receipt shows on it in the Splitwise app. Pairs with the scan flow: extract line-items from the image, then attach the image itself.

  • Statement import — import_statement(transactions, default_split_name?, dedup) proposes a categorized, split-suggested, duplicate-flagged list from parsed statement rows; confirm_import(rows) bulk-creates the approved ones via the itemization engine.

  • Gmail read-only connector (optional) — gmail_find_statements / gmail_read_statement fetch bank/card statement emails (scope gmail.readonly) to feed statement import — the CRED-style "no manual entry" flow. Requires a one-time Google Cloud OAuth setup; install extras with pip install -e ".[gmail]". The connector only reads email text; all expense creation still goes through the reviewed import path.

  • Save default splits — save_default_split(name, split) / list_default_splits / delete_default_split. Reuse a template by putting "split_ref": "roomies-4way" on an item. Stored locally in ~/.expensifyai/splits.json.

Pick any date range — the whole dashboard recomputes live in the browser:

Filtered to one month

Try it without an account:

python examples/generate_demo.py   # writes examples/demo-dashboard.html

Features (MCP)

  • Full API Access: Manage expenses, groups, friends, and comments.

  • Natural Language Resolution: Fuzzy matching for names ("John" -> "John Smith") and groups.

  • Dual Auth: Supports both OAuth 2.0 (recommended) and API Keys.

  • Smart Caching: Optimizes performance for static data like categories and currencies.

Prerequisite: Splitwise Pro

Since early September 2026 the Splitwise API requires the API app's registered developer to hold an active Splitwise Pro subscription (~$3/mo, $30/yr). Without it every call returns:

403 — base: This app can't access Splitwise right now because its developer
      doesn't have an active Splitwise Pro subscription.

This is Splitwise's policy, not a bug here — their API terms reserve it explicitly ("Splitwise may impose conditions on the use of the Self-Serve API, including, for example, maintaining an active Splitwise Pro subscription"). Verified: the same 403 is returned by this server, by a bare API client with no server involved, and to keys that worked days earlier. There is no workaround — the gate sits upstream of every endpoint, reads included.

So to use the 37 Splitwise tools you need Pro, then register an app at secure.splitwise.com/apps for your key. The 5 statement tools (Gmail, PDF unlock, bank store) need no Splitwise account at all — 3 of them work entirely offline.

Installation

git clone https://github.com/udaysrinu/ExpensifyAI
cd ExpensifyAI
python -m venv venv
source venv/bin/activate
pip install -e .

Configuration

See SETUP.md for detailed authentication and configuration instructions.

Quick Config

Run the included setup script:

python -m splitwise_mcp_server.oauth_setup

Use the keys provided there, and add all three to your mcp.json:

{
  "mcpServers": {
    "splitwise": {
      "command": "python",
      "args": ["-m", "splitwise_mcp_server"],
      "env": {
        "SPLITWISE_OAUTH_ACCESS_TOKEN": "your_token_here"
      }
    }
  }
}

Get your Auth Keys You can get your Consumer Key and Secret by registering an app at https://secure.splitwise.com/apps.

IMPORTANT: Using a Virtual Environment? If you installed the package in a venv or Conda environment, you must use the absolute path to the python executable in your config.

"command": "/absolute/path/to/venv/bin/python"

See SETUP.md for details.

Usage

The server enables natural language interactions with your Splitwise data.

Examples:

  • "What's my current balance?"

  • "Split a $50 dinner with Sarah."

  • "Use the receipt I uploaded to split the dinner between Manav and me."

  • "Show me expenses from last month."

  • "Create a group called 'Ski Trip' with Mike."

Hosted mode — use it from your phone, with your own key

The stdio server above runs as a subprocess on your machine, which the Claude mobile and web apps cannot do. api/index.py is the same server over HTTP, deployable to Vercel and addable as a custom connector.

You bring your own key; the server stores nothing. Add the connector with your Splitwise API key (from secure.splitwise.com/apps) as a bearer token:

URL:    https://<your-deployment>.vercel.app
Header: Authorization: Bearer <your Splitwise API key>

Each request is served with the credential it arrives with, so one deployment works for any number of people, and the operator never holds anyone else's key. Do not set SPLITWISE_API_KEY on a shared deployment — it becomes the fallback for keyless requests, which turns the public URL into your own account. See per_request.py.

The hosted server has 37 of the 42 tools. Left out: gmail_find_statements, gmail_read_statement, unlock_statement_pdf, import_statement, confirm_import. Those need your Gmail token, your statement PDFs and the SQLite store under ~/.expensifyai. A serverless filesystem is ephemeral so they could not work anyway — and hosting them would put your bank statements and mailbox credentials on a server, which defeats the point of keeping that data local. Bank statement import stays on stdio.

vercel            # deploy; no environment variables to set

Because a key is tied to the app that issued it, bring-your-own-key means each user registers their own app — and so each needs their own Pro (see Prerequisite). If you would rather carry that cost yourself so your users pay nothing, switch this to OAuth2: one app you register, your Pro, and users just authorize. OAuth2Handler already emits the same Authorization: Bearer <token> header, so per_request.py needs no changes — only the token's source does. oauth_setup.py has the authorization-code flow.

Tools

See TOOLS.md for detailed documentation.

User Tools

  • get-current-user: Get authenticated user information

  • get-user: Get information about a specific user

Expense Tools

  • create-expense: Create a new expense with splits

  • get-expenses: List expenses with optional filters

  • get-expense: Get detailed expense information

  • update-expense: Update an existing expense

  • delete-expense: Delete an expense

Group Tools

  • get-groups: List all groups

  • get-group: Get detailed group information

  • create-group: Create a new group

  • delete-group: Delete a group

  • add-user-to-group: Add a user to a group

  • remove-user-from-group: Remove a user from a group

Friend Tools

  • get-friends: List all friends

  • get-friend: Get detailed friend information

Resolution Tools

  • resolve-friend: Fuzzy match friend names to user IDs

  • resolve-group: Fuzzy match group names to group IDs

  • resolve-category: Fuzzy match category names to category IDs

Comment Tools

  • create-comment: Add a comment to an expense

  • get-comments: Get all comments for an expense

  • delete-comment: Delete a comment

Utility Tools

  • get-categories: Get all expense categories

  • get-currencies: Get all supported currencies

Arithmetic Tools

  • add: Add multiple numbers

  • subtract: Subtract numbers

  • multiply: Multiply numbers

  • divide: Divide numbers

  • modulo: Calculate remainder

Development

# Setup
git clone https://github.com/udaysrinu/ExpensifyAI
cd ExpensifyAI
python -m venv venv
source venv/bin/activate
pip install -e ".[dev]"

# Test
pytest                              # full suite
pytest tests/test_analytics.py      # deterministic analytics (15 tests)

# See the dashboard with no account
python examples/generate_demo.py    # writes examples/demo-dashboard.html

License

MIT License. See LICENSE for details.

Built on top of tarunn2799/splitwise-mcp (MIT); the analytics engine, interactive dashboard, and tests are added by ExpensifyAI.

Available Tools

42 tools
add_user_to_groupAdd User To GroupB

Add a user to a group by user_id or by email (with first_name/last_name for new invites).

ParametersJSON Schema
NameRequiredDescriptionDefault
emailNo
user_idNo
group_idYes
last_nameNo
first_nameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description alone must disclose behavioral traits. It indicates a mutation (adding a user) but does not state permissions, idempotency, whether invites are sent, what happens if the user already belongs to the group, or the effect of the email-only path.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence that front-loads the action and packs in the two identification methods and the invite condition. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the core action and parameter conditions, and an output schema exists for return values. However, with no annotations and no side-effect/error behavior, the definition is not fully complete for a mutating tool; edge cases like duplicate membership or missing identifiers are unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 0%, the description carries the burden. It explains the roles of user_id, email, first_name, and last_name (the latter two for new invites). It leaves group_id implied by context, and it doesn't explicitly state that one of user_id/email is required, but the 'by...or by...' construction conveys this.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Add') and resource ('a user to a group'), with the two identification modes. It is distinguishable from its sibling remove_user_from_group by the opposite action, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for new invites' implies a conditional for using email plus names, which is a light usage cue. But the description offers no explicit guidance on when to use this tool over alternatives like remove_user_from_group or create_user, nor any exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

analyze_spendingAnalyze SpendingA

Deterministic spending analytics for the current user, a group, or a friend.

All numbers are computed in Python (never estimated by the model): category breakdown, monthly trend, owed-vs-paid ("mine vs split"), transaction ledger, top transactions, and — for groups — per-member comparison, category×member matrix, and a minimum-transaction settlement plan. Every result includes a reconciliation check (shares must sum to cost) and a multi-currency guard.

target_type: "me" (all your expenses), "group" (needs target_id=group_id), or "friend" (needs target_id=friend user_id). dated_after / dated_before: ISO 8601 date filters (optional). generate_dashboard: if True, also writes a self-contained HTML dashboard and returns its file path in dashboard_path. Perspective is always the authenticated user.

ParametersJSON Schema
NameRequiredDescriptionDefault
top_nNo
target_idNo
dated_afterNo
target_typeNome
dated_beforeNo
generate_dashboardNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses deterministic computation (not model estimation), reconciliation checks, a multi-currency guard, the authenticated-user perspective, and the side effect of generating an HTML dashboard. This is substantial transparency, though it does not mention permissions or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a bit long but well-organized: a clear opening sentence, a detailed output list, then parameter explanations. It is front-loaded with the core purpose, and each sentence adds value without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 6 optional parameters, an output schema, and no annotations, the description covers the essential behavioral traits and parameter semantics. It does not explain the output schema structure, but that is provided separately. It is complete for an agent to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must explain parameters. It explicitly describes target_type, target_id, date filters (ISO 8601), and generate_dashboard's behavior and output path. top_n is implied via 'top transactions'. All parameters are effectively covered, exceeding the compensation expected at 0% schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool computes deterministic spending analytics for the user, group, or friend, and lists specific outputs like category breakdown, monthly trend, and settlement plans. It distinguishes itself from raw data retrieval siblings like get_expenses by emphasizing computed analytics rather than simple listing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains target_type and when to use each scope (me, group, friend) and implies this tool is for analytics rather than raw data. It does not explicitly name alternatives or exclusions, but the context is clear enough for an agent to select it when computed insights are needed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

attach_receiptAttach ReceiptA

Attach a receipt image or PDF (from a local file path) to an existing expense.

Uploads the file to Splitwise via multipart on update_expense. Supports common image types and PDF. The receipt then shows on the expense in the app/website. (Splitwise's API only exposes receipt UPLOAD, not OCR — pair this with create_itemized_expense, where the calling agent reads the image for line-items.)

ParametersJSON Schema
NameRequiredDescriptionDefault
expense_idYes
image_pathYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the tool uploads via multipart on update_expense (which implies a write operation) and notes that the API does not support OCR. However, it doesn't mention potential side effects like requiring existing expense ownership or failure modes, but the core behavior is reasonably clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: a clear first sentence stating the action and inputs, followed by implementation details, and a note about OCR pairing. All sentences provide value, no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with 2 parameters, and the description covers the key inputs, file type support, and the important caveat about lack of OCR. An output schema exists, so return values are documented elsewhere. A minor gap is the lack of explicit error handling or authorization requirements, but for the tool's simplicity, it's largely complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has only parameter names and types (expense_id: integer, image_path: string) with 0% coverage. The description explains that image_path is a local file path and mentions supported file types (image and PDF), adding some meaning beyond the schema. However, it doesn't specify formats or constraints for expense_id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: attach a receipt image or PDF from a local file path to an existing expense. It specifies the verb (attach), resource (receipt to expense), and input types (image or PDF). This distinguishes it from other expense-related tools in the sibling list such as create_expense and update_expense.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is used when uploading a receipt to an existing expense and mentions pairing with create_itemized_expense for OCR. However, it doesn't explicitly state when not to use this tool or mention alternatives for other receipt-related tasks.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare_group_membersCompare Group MembersA

Deterministic per-member comparison for a group: total spend + ranking, category×member matrix, insights (highest/lowest/average/spread), and a minimum-transaction settlement plan. Convenience wrapper over analyze_spending with target_type='group'.

ParametersJSON Schema
NameRequiredDescriptionDefault
group_idYes
dated_afterNo
dated_beforeNo
generate_dashboardNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It communicates determinism, enumerates the analysis dimensions, and explains that it delegates to analyze_spending, making the read-only analysis intent reasonably clear. It does not explicitly state that no data is mutated, but the analytic language strongly implies it.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core behavior and output list, with a succinct wrapper note. Every word contributes to understanding the tool's function and relationship to a sibling.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description gives a clear high-level behavior and output set, and an output schema exists for return values. However, with no annotations and zero schema coverage, the optional date filters and dashboard flag are left undocumented, making the description incomplete for fully utilizing the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for parameter meaning. It only clarifies that group_id identifies the group being compared; dated_after, dated_before, and generate_dashboard are entirely unexplained, leaving an agent unable to correctly use the optional parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: deterministic per-member comparison for a group. It enumerates the concrete outputs (total spend, ranking, category×member matrix, insights, settlement plan) and explicitly distinguishes itself from analyze_spending as a convenience wrapper with target_type='group'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly indicates when to use the tool: for per-member group comparison, framed as a convenience wrapper over analyze_spending. It identifies the relevant alternative but does not explicitly state conditions for not using it or other alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

confirm_importConfirm ImportA

Bulk-create expenses from approved statement rows (after import_statement review).

rows: approved items, each {date, description, amount, category?, split_ref?}. Rows with a split_ref use that saved template; otherwise 100%-personal. group_id: default group for created expenses (0 = non-group). Per-row group_id overrides. Each expense is created via the itemization engine (exact paise, reconciled). Returns a per-row result list (created id or error). Sequential to stay under rate limits.

ParametersJSON Schema
NameRequiredDescriptionDefault
rowsYes
group_idNo
currency_codeNoINR

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and handles it well: it discloses sequential execution to stay under rate limits, per-row result/error returns, exact-paise/reconciled itemization engine behavior, and split_ref fallback logic. No annotation contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: purpose first, then parameter semantics, then execution and return behavior. Every sentence earns its place, with no filler or repetition of schema defaults.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers input shape, defaults, per-row behavior, return format, and the required prior workflow step, which is enough for correct invocation. The main gap is currency_code semantics and slightly more explicit prerequisites for what makes rows 'approved', but the overall picture is strong.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides 0% description coverage, but the description compensates by documenting the rows item shape and group_id behavior including the 0=non-group default and per-row override. However, currency_code is not described at all, leaving one of the three parameters to inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Bulk-create expenses from approved statement rows (after import_statement review)', giving a specific verb, resource, and workflow stage. This clearly differentiates it from single-expense creation via create_expense and from the parsing/review step of import_statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly places the tool 'after import_statement review', giving a clear precondition and context for use. It does not name alternatives or say when not to use it, but the workflow placement is enough to guide an agent to the correct stage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_commentCreate CommentB

Add a comment to an expense. Visible to all users in the expense.

ParametersJSON Schema
NameRequiredDescriptionDefault
contentYes
expense_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It does add one useful behavioral trait ('Visible to all users'), but it omits other side effects like whether a notification is triggered, whether the expense is modified, or whether there are any permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no fluff. The action is front-loaded ('Add a comment to an expense'), and the second sentence adds a useful visibility detail without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (2 params, output schema present), so the description is close to adequate. However, it lacks usage guidance and parameter clarification, which an agent needs to decide between this tool and its siblings. The visibility detail adds some completeness but does not cover prerequisites or alternatives.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description should compensate by explaining parameters. It only refers to 'an expense' generically and never names content or expense_id, leaving the agent to infer that content is the comment text and expense_id is the target expense. The parameter names are somewhat self-explanatory but the description adds no explicit mapping.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Add a comment') and the resource ('an expense'), and adds the scope 'Visible to all users in the expense.' It does not explicitly name sibling tools like get_comments or delete_comment, but the verb makes the distinction evident, so it misses the full 5 for sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to use this tool vs get_comments (to read comments) or delete_comment (to remove them). There are also no prerequisites mentioned, such as requiring the user to be a member of the expense or that the expense must exist.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_expenseCreate ExpenseA

Create a new expense. Cost is a string with 2 decimals (e.g. "25.50"). Splits equally by default; provide users list with paid_share/owed_share for custom splits. Each user needs user_id or (email + first_name + last_name). Set repeat_interval to "weekly", "fortnightly", "monthly", or "yearly" for recurring expenses.

ParametersJSON Schema
NameRequiredDescriptionDefault
costYes
dateNo
usersNo
detailsNo
group_idNo
category_idNo
descriptionYes
currency_codeNoUSD
split_equallyNo
repeat_intervalNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral disclosure burden. It discloses meaningful behaviors: cost must be a string with 2 decimals, equal splitting is the default, custom split requires paid_share/owed_share, and user identity must be user_id or a full email/first/last combination. These details go beyond the schema and clarify how the API behaves.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured. It front-loads the core purpose, then gives the most important behavioral requirements in three additional dense sentences. Every sentence contributes necessary information without fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (10 parameters, no annotations) the description covers the most error-prone areas: cost format, custom split inputs, user identity, and recurrence values. An output schema exists, so return value details are not required. It is nearly complete for a typical agent invocation, though it does not clarify how it differs from create_itemized_expense or explain the group_id default behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds essential semantics for cost formatting, the structure of the users array, and valid repeat_interval values. Some parameters like date, currency_code, group_id, and category_id are not explained in the description, though their names are fairly self-explanatory in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action and resource: 'Create a new expense.' It adds meaningful context about cost formatting, splits, and recurrence, making the tool's purpose obvious. However, it does not distinguish itself from the sibling create_itemized_expense, so it misses full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides conditional usage guidance: use custom splits when paid_share/owed_share are provided, and set repeat_interval for recurring expenses. It does not explicitly state when to prefer this tool over create_itemized_expense or other sibling tools, and it offers no exclusion criteria. This is better than no guidance, but not fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_friendCreate FriendB

Add a friend by email address. Optionally provide their first and last name.

ParametersJSON Schema
NameRequiredDescriptionDefault
user_emailYes
user_last_nameNo
user_first_nameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only states the basic create action and does not mention side effects, permission requirements, duplicate handling, or what happens when the email address does not correspond to an existing user.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler. The core action and required parameter are front-loaded, and the optional parameters are stated in a single follow-up clause.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple, an output schema exists, and the description covers the required and optional parameters, so an agent can make a basic call. However, it lacks any note about duplicate handling, whether the friend must already be a platform user, or what response to expect, leaving the agent to discover these behaviors at runtime.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description must compensate, and it does name all three parameters: user_email is implied by 'by email address,' and user_first_name/user_last_name are referenced by 'optionally provide their first and last name.' However, it adds no format details, constraints, or clarification beyond the field names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Add') and resource ('a friend by email address'), making the operation unambiguous. It clearly distinguishes this from sibling tools like get_friends, delete_friend, and resolve_friend by stating the action is creation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool versus alternatives like resolve_friend or get_friends. It states the mechanism (add by email) but provides no conditions, exclusions, or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_groupCreate GroupA

Create a new group. group_type is one of: home, trip, couple, other. Optionally add initial members via users list with user_id or email+name.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
usersNo
group_typeNoother
simplify_by_defaultNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of disclosing side effects. It states that it creates a group and optionally adds members, but it does not mention permissions, whether the operation is reversible, what happens if invalid user data is supplied, or what simplify_by_default does beyond its schema default.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two compact sentences with no filler. The core action is front-loaded, and the parameter guidance is presented in a readable, scannable format.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The definition is adequate for a simple creation tool, and the presence of an output schema reduces the need to describe return values. But with no annotations and 0% schema description coverage, the unexplained simplify_by_default parameter and absence of sibling-contrast guidance leave meaningful gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds explicit allowed values for group_type and the accepted user identifier shapes (user_id or email+name), which is genuinely helpful. However, it omits simplify_by_default entirely and does not clarify the required name field beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a direct verb and resource: 'Create a new group.' It further clarifies scope by listing valid group_type values and the optional initial-members behavior, which clearly distinguishes it from sibling tools like delete_group, get_group, and add_user_to_group.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit statement of when to use this tool versus alternatives. The phrase 'Optionally add initial members' implies that creation-time member addition belongs here rather than in add_user_to_group, but it never names a sibling or gives a when-not-to-use condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_itemized_expenseCreate Itemized ExpenseA

Create ONE Splitwise expense from itemized line-items, each with its OWN split.

This is how a receipt becomes an expense: the agent extracts line-items from the receipt image and passes them here. Each item can split differently (e.g. beers 3/4 to one person, groceries 4-way, cake between two) — the tool computes each person's total owed_share exactly in integer paise and reconciles to the total before writing. Set dry_run=True to preview the computed split without creating.

items: list of { "desc": "Beers", "amount": "2710.00", # rupees, string with 2 decimals "category": "Drinks", # optional (free text, informational) "paid_by": , # who fronted this item "split": { # OR "split_ref": "" "type": "equal" | "shares" | "exact", "among": [user_id, ...], # for equal/shares "shares": {user_id: weight, ...}, # for shares "exact": {user_id: "amount", ...} # for exact (must sum to amount) } }

ParametersJSON Schema
NameRequiredDescriptionDefault
dateNo
itemsYes
dry_runNo
group_idYes
descriptionYes
currency_codeNoINR

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does well: it reveals that the tool computes each person's owed share in integer paise, reconciles to the total before writing, supports per-item split modes, and that dry_run avoids creating anything. This goes well beyond the bare input schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but the embedded items structure is the clearest way to convey an inherently complex parameter. It is front-loaded with the core purpose and uses the inline JSON efficiently, though a few phrases could be trimmed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with a free-form items schema, the description gives enough detail to construct valid calls, covering split modes, amount formats, and dry_run behavior. It does not discuss date formats or error conditions, but the output schema likely covers the return shape and the required fields are obvious.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description compensates by fully documenting the complex items parameter, including amount formatting, split types, paid_by, and split_ref. It also explains dry_run. It does not describe date, currency_code, group_id, or description, though those are relatively self-explanatory.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Create ONE Splitwise expense') on a specific resource, and qualifies it as being built from itemized line-items each with its own split. This clearly distinguishes it from the sibling create_expense tool, which handles a single split.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context: 'This is how a receipt becomes an expense' and explains that line-items are extracted from a receipt image and passed here. It also mentions dry_run for preview, but it does not explicitly state when to prefer this over create_expense or list exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_commentDelete CommentA

Delete a comment. You can only delete your own comments.

ParametersJSON Schema
NameRequiredDescriptionDefault
comment_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden of behavioral disclosure. It adds the ownership restriction but omits whether deletion is permanent or irreversible, whether replies or related data are cascade-deleted, and what side effects may occur.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler; the action is stated first and the key restriction follows immediately. This is appropriately sized for a one-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter destructive operation with an output schema, the description covers the target resource and authorization requirement. It is incomplete only in that it does not explicitly state destructive consequences or reversibility.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not explain comment_id or how it should be obtained, so it fails to compensate for the schema gap. The parameter name and integer type are self-explanatory, but the description adds no semantic value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses a specific verb-resource pair ('Delete a comment') and unambiguously distinguishes itself from sibling tools like create_comment and get_comments. The ownership constraint further clarifies the operation's exact scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states a condition of use: 'You can only delete your own comments,' which tells an agent when the tool is not applicable. It does not name an alternative, but no sibling tool performs comment deletion, so the guidance is effectively sufficient for selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_default_splitDelete Default SplitB

Delete a saved split template by name.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It makes the destructive nature clear via 'Delete' and identifies what is removed, but it does not state whether deletion is permanent or irreversible, what happens if the name does not exist, or whether deleting a template affects existing splits. This is a significant gap for a destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no filler or redundancy. Every word earns its place, and the key information (action, object, locator) is compactly delivered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple one-parameter schema and the presence of an output schema, the description is minimally viable: an agent knows what to call and what argument to pass. However, it lacks behavioral caveats and usage context that would be valuable for a destructive tool with no annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does so minimally by explaining that the sole required parameter 'name' is the identifier of the template to delete. It does not add format, constraints, or examples, but for a simple string identifier this is adequate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('Delete'), the resource ('saved split template'), and how it is identified ('by name'). It is unambiguous about what the tool does, though it does not explicitly contrast itself with sibling tools like delete_group or save_default_split.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the imperative 'Delete' and the resource type: an agent can infer this tool is for removing a saved split template, while save_default_split and list_default_splits serve complementary purposes. However, no explicit guidance is provided about prerequisites, when not to use this tool, or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_expenseDelete ExpenseA

Delete an expense permanently. Use restore_expense to undo.

ParametersJSON Schema
NameRequiredDescriptionDefault
expense_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description has the full burden and does disclose that deletion is permanent and that restore_expense can undo it. However, the wording is slightly internally inconsistent ('permanently' vs. undo), and it does not mention side effects on related data such as receipts, comments, or group balances.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the key fact ('permanently') and ending with the actionable undo path. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter destructive tool, the description covers the core behavior and undo path, and an output schema exists so return values need not be described. The lack of detail about what actually gets destroyed (only the expense row, or associated data) and the permanent/undo ambiguity leave an agent with meaningful uncertainty.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for the expense_id parameter, but it does not mention it. The parameter name is self-explanatory, but the description adds no semantic value beyond the schema field itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('Delete an expense permanently') on a clear resource, and the permanence qualifier sets expectations. It is readily distinguished from sibling tools like delete_group or delete_comment by naming the resource type.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly names restore_expense as the way to undo the operation, giving an agent a clear remediation path. It does not list when to avoid deleting vs updating, but for a simple deletion the context is adequate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_friendDelete FriendA

Remove a friendship. Does not affect shared expenses or balances.

ParametersJSON Schema
NameRequiredDescriptionDefault
friend_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds a meaningful non-effect ('Does not affect shared expenses or balances'), which is useful context for a destructive operation. However, it does not mention reversibility, permissions, or any cascading effects beyond the relationship itself.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler. The core action is front-loaded, and the important caveat about expenses and balances is stated directly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter delete operation with an output schema, the description is largely complete: it states the action and a key non-obvious side-effect. It could add more about the semantics of friend_id, but the tool is simple enough that an agent can infer correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description should compensate by explaining what friend_id means and how it is used. It does not mention the parameter at all, though the property name 'friend_id' and the phrase 'Remove a friendship' provide some implicit meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Remove') and a specific resource ('a friendship'), making the tool's purpose immediately clear. The phrasing also distinguishes it from sibling tools like delete_group by targeting the friendship relationship rather than a group.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool versus alternatives such as remove_user_from_group or create_friend. The purpose is implied by the name and first sentence, but there is no stated condition, exclusion, or alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_groupDelete GroupA

Delete a group. All expenses must be settled first.

ParametersJSON Schema
NameRequiredDescriptionDefault
group_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It discloses the precondition but does not mention permanence, cascading effects, or error handling. The precondition is useful, but the description is thin on broader behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The action is front-loaded, and the precondition is stated immediately after. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple delete tool with one obvious parameter and a clear precondition, the description is nearly complete. It lacks return-value or error details, but the output schema exists (though not shown) and the simplicity reduces the need for more.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does not mention group_id at all, nor explain its role. The schema only provides type and name, so the agent must infer meaning from the parameter name alone. This is a notable gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (delete) and the resource (group), making it distinct from sibling tools like delete_expense or delete_friend. It also adds a critical precondition, which further clarifies the purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The precondition 'All expenses must be settled first' provides clear usage guidance, implying when not to use the tool. It doesn't name alternatives, but the constraint is explicit and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_categoriesGet CategoriesA

Get all expense categories and subcategories. Results are cached.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It adds the useful trait that 'Results are cached,' implying possible staleness, and 'Get' implies a read-only operation. However, it does not mention permissions, potential side effects, or how caching might affect freshness.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler. The key action and scope are front-loaded, and the caching note is concise and relevant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only tool with an output schema, the description is mostly sufficient. It states the resource and a behavioral caveat (caching). It could be slightly stronger with a note about when the cached data refreshes, but it is otherwise complete enough for effective tool invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the baseline is 4. The description does not need to explain parameter semantics and adds no confusing or contradictory parameter information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb and resource: 'Get all expense categories and subcategories.' It is unambiguous about scope. However, it does not explicitly distinguish itself from the sibling tool 'resolve_category', so it misses the top score for sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool versus alternatives like resolve_category. The description only states what the tool does, not the appropriate usage context or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_commentsGet CommentsA

Get all comments on an expense.

ParametersJSON Schema
NameRequiredDescriptionDefault
expense_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It clearly indicates a read-only operation, but does not describe ordering, pagination, error behavior, or what happens when the expense does not exist. Minimal but not misleading.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, direct sentence with no filler. The core action and target are front-loaded, making it easy for an agent to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple getter with one required parameter and an output schema, the description is largely sufficient. It lacks some behavioral context, but the output schema covers return shape and the input schema covers required parameters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It only loosely maps expense_id to 'an expense' without explaining the parameter's meaning, constraints, or expected format beyond what the property name already implies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Get all comments on an expense.' This clearly distinguishes it from sibling tools like create_comment and delete_comment, which perform different operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when a user needs all comments tied to an expense, but it does not explicitly state when to use this over alternatives or mention any exclusions. No alternative comment-retrieval tool exists, so the implication is adequate but not explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_currenciesGet CurrenciesA

Get all supported currency codes and symbols. Results are cached.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It adds a genuine non-obvious trait: 'Results are cached,' which informs the agent that data may be stale or not freshly fetched. It does not mention cache duration, but for a simple no-param read tool this is sufficient.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, both informative: the first states the purpose, the second discloses caching behavior. There is no filler, repetition, or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read-only lookup with an output schema present, the description is complete. It states what the tool returns and the one relevant behavioral nuance (caching), and it does not need to explain return values because the output schema covers that.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema is vacuous and there is nothing to document. The description adds value by clarifying exactly what data is returned (currency codes and symbols), which is the relevant semantic content for this tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Get'), a precise resource ('all supported currency codes and symbols'), and this clearly distinguishes it from the sibling tools, none of which overlap with a currency-list lookup. There is no ambiguity about what the tool returns.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, exclusions, or named alternatives. However, the tool is a zero-parameter getter for a unique domain, so the intended use is reasonably implied. Still, the description leaves the agent to infer context rather than stating it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_current_userGet Current UserA

Get the current authenticated user's profile (id, name, email, picture).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Get' strongly implies a read-only operation with no side effects, and listing profile fields is helpful, but the description does not explicitly state read-only behavior, auth requirements beyond 'authenticated,' or any other behavioral traits. Still, for a simple getter this level is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that communicates purpose and output fields with zero wasted words. It is optimally concise for a simple no-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no parameters gagliardetto and an output schema exists, the description fully covers what an agent needs: what the tool does, whose profile it returns, and which fields are included. No significant information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there are no parameter semantics to document. The baseline for a zero-parameter tool is 4, and the description adds value by specifying the returned profile fields, which is sufficient given the empty input schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), a clear resource ('current authenticated user's profile'), and enumerates the returned fields (id, name, email, picture). This distinguishes it from sibling get_user by emphasizing the current authenticated user, and it is immediately clear what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'current authenticated user' provides clear context for when to use the tool: when you need the profile of the logged-in user. However, it does not explicitly contrast with sibling get_user or state when not to use this tool, leaving the comparison to the agent's inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_expenseGet ExpenseA

Get full details of a single expense including users, splits, and comments.

ParametersJSON Schema
NameRequiredDescriptionDefault
expense_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states only that the tool retrieves details; it does not mention read-only semantics, error behavior for missing expense_id, authentication requirements, or any side effects. The term 'Get' implies a read operation, but that inference is not explicitly confirmed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, tightly worded sentence with no filler. It front-loads the action and resource, then specifies the included detail types, all in 12 words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter getter with an output schema, the description is workable but leaves gaps. It does not address error handling, relationship to sibling tools like get_expenses or get_comments, or any prerequisites. Given the absence of annotations, this is a minimum-viable description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate by explaining expense_id. Instead, it only mentions output facets (users, splits, comments) and provides no guidance on how to obtain a valid expense_id, its format, or constraints. The description adds no meaning beyond the schema's bare type declaration.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses the specific verb 'Get' with the resource 'expense' and explicitly says 'single expense', which immediately distinguishes it from the sibling get_expenses. It also enumerates what 'full details' includes (users, splits, comments), leaving no ambiguity about the tool's purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'full details of a single expense' provides a clear context for when to use this tool: whenever you need comprehensive information about one expense. It does not, however, explicitly state when not to use it or name alternatives like get_expenses, so it falls short of full exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_expensesGet ExpensesC

List expenses with optional filters. Dates are ISO 8601 format. Max limit is 100.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
group_idNo
friend_idNo
dated_afterNo
dated_beforeNo
updated_afterNo
updated_beforeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral disclosure burden. It does disclose the 100-result maximum and ISO 8601 date format, but it omits key behaviors such as how filters combine, whether group_id and friend_id are mutually exclusive, default sort order, and pagination semantics. Partial transparency at best.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, followed by the two most important invocation constraints (date format and result cap). Every sentence earns its place, though for 8 undocumented parameters, slightly more detail would have improved the tradeoff without hurting readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite the output schema covering return values, the tool has 8 parameters with zero schema descriptions and zero annotations. The description only addresses date format and limit cap, leaving filter semantics, parameter relationships, and pagination behavior unspecified. For a list endpoint of this complexity, the definition is incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all 8 parameters. It adds meaning to the date parameters (ISO 8601) and limit (max 100), but says nothing about group_id, friend_id, or offset beyond calling everything 'optional filters.' The compensation is only partial for a large, fully undocumented parameter set.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource ('List expenses') and indicates the tool is a filtered listing operation. It is clear enough to be distinguished from singular get_expense, but it does not explicitly differentiate from search_expenses or other listing siblings, and the name itself carries the differentiation load.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description offers no guidance on when to use this tool versus get_expense, search_expenses, or analyze_spending. It states that filters exist but provides no context on typical use cases, prerequisites, or exclusions, leaving the agent to infer selection criteria from names alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_friendGet FriendB

Get a friend's details including balances and shared groups.

ParametersJSON Schema
NameRequiredDescriptionDefault
user_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that returned details include balances and shared groups. However, it does not state side effects, failure behavior, or prerequisites such as whether the friendship must already exist, leaving a visible but moderate gap for a simple getter.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The whole description is one front-loaded sentence containing the action, resource, and key return aspects. There is no repetition, filler, or restating of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, low-complexity tool with an output schema, the description is nearly adequate: it previews the main return areas and the input is inferable. It is not fully complete because it omits any relation to sibling retrieval tools and does not clarify the meaning of user_id, but the tool's simplicity and output schema mitigate the omissions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the tool description does not define user_id or explicitly say it is the friend's ID. The parameter name is suggestive, but the description adds no semantics beyond the schema's type and required flag, so it fails to compensate for the missing schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb 'Get' and names a specific resource, 'a friend's details,' further specifying 'including balances and shared groups.' This makes the tool's scope clear and helps separate it from sibling tools like get_friends and get_user, though it never explicitly contrasts them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to choose this tool over get_friends, resolve_friend, or get_user. The description only implies a single use case by naming the operation; it does not provide context, exclusions, or alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_friendsGet FriendsA

List all friends with their balance information.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'List' implies a read-only operation and 'balance information' hints at the output, but it does not disclose pagination, ordering, scope (current user's friends vs all system friends), or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler. The action and output focus are front-loaded and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-parameter, read-only list tool with an output schema, the description covers the essential selection and invocation information. It would benefit from a pointer to get_friend for singular lookups, but that absence is not crippling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so the baseline is 4. There are no parameter semantics for the description to clarify.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List'), a clear resource ('friends'), and the returned data ('balance information'). It distinguishes itself from the sibling get_friend by covering the plural/all-friends case.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus get_friend or other friend-related tools. The word 'all' weakly implies bulk listing, but no alternatives or exclusions are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_groupGet GroupA

Get a group's details including members, balances, and simplified debts.

ParametersJSON Schema
NameRequiredDescriptionDefault
group_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of conveying behavior. The verb 'Get' clearly indicates a read-only operation, and the description indicates what is returned. However, it does not mention permissions, error behavior, or any side-effect guarantees.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that immediately states the operation and result scope. Every word earns its place, with no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with one required parameter and an output schema, so the description does not need to enumerate return fields. It names the core content (members, balances, simplified debts) and is nearly complete; only explicit usage differentiation and auth/error context are absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the single parameter group_id is self-explanatory as an integer identifier. The description's phrase 'a group's details' reinforces the parameter's role without explicitly defining it. This is adequate but adds little beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get'), a specific resource ('a group's details'), and the key content included (members, balances, simplified debts). The singular 'a group' clearly distinguishes this from the sibling get_groups tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when a single group's details are needed, but it gives no explicit guidance about when to use this tool instead of get_groups or other siblings. No exclusions or alternative routing are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_groupsGet GroupsA

List all groups the current user belongs to, with members and balances.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of behavioral disclosure. It does communicate a read-only list operation and user scoping, but it omits details like empty-result behavior, errors, pagination, or whether balances are computed or stored.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no wasted words. It states the action, scope, and returned data efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple zero-parameter list tool with an output schema, the description covers the essential information: what is listed, for whom, and what is included. It could be slightly more complete by referencing get_group for single-group lookups, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and schema coverage is 100%, so the baseline for this dimension is 4. The description adds useful context about the returned groups but does not need to explain parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('List') with a clear resource ('all groups') and scope ('the current user belongs to'), and includes return content ('with members and balances'). This clearly distinguishes it from the singular sibling tool get_group.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context: use it to list all groups for the current user, which implies the typical use case. However, it does not explicitly name alternatives like get_group for retrieving a single group or state when not to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_notificationsGet NotificationsA

Get recent notifications for the current user (new expenses, payments, comments, group activity).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It explains scope but does not state that this is a read-only operation, whether authentication is required, or any limits on recency or result count. The categories add some context, but the safety profile is left implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that conveys the tool's purpose and scope with no wasted words. Every element adds useful information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read with an output schema present, the description covers the essential selection context: current user, recency, and event types. The only minor gap is the unquantified meaning of 'recent', but this does not block correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the baseline is 4. The description appropriately avoids inventing parameter details and focuses on the output scope.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Get'), a specific resource ('notifications'), a scope ('current user'), and a timeframe ('recent'). It also lists the kinds of events covered, which clearly separates it from siblings like get_comments or get_expenses.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes clear this returns the current user's recent notifications, which implies when to use it. There is no obvious sibling alternative for notifications, so explicit when-not-to-use guidance is less necessary; the context is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_userGet UserB

Get a user's profile by their ID.

ParametersJSON Schema
NameRequiredDescriptionDefault
user_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It indicates a read-only 'get' operation, but it does not disclose any potential errors, required authorization, or behavior for missing IDs. This is minimal but not misleading; it adequately implies a simple fetch without going into edge cases. Given the simplicity, a 3 is appropriate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no redundancy. It front-loads the core action ('Get') and immediately specifies the resource and criterion. Every word earns its place, making it an exemplar of conciseness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter, no nested objects, and an existing output schema), the description is sufficiently complete. It does not need to detail return values because the output schema covers that. The only minor gap is the lack of guidance on the distinction from get_current_user, but as a standalone description, it provides enough for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema lists only 'user_id' with an integer type and zero description coverage. The description adds meaning by stating that the profile is retrieved 'by their ID', which clearly implies that the parameter is the user's unique identifier. However, it does not explicitly label it as required or explain its format beyond the type. This partial clarification justifies a 3, falling slightly short of fully compensating for the schema's lack of parameter descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb ('Get'), a specific resource ('user's profile'), and the method of selection ('by their ID'). It is precise and unambiguous, but it does not explicitly distinguish itself from the sibling tool get_current_user, which serves a related but different purpose. The phrasing implies the difference, but without an explicit note, it earns a 4 rather than a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus the alternatives. Notably, get_current_user exists as a sibling and would be the natural choice for the authenticated user, but the description does not mention it or any context that should trigger this tool. It simply states the operation with no when/when-not instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gmail_find_statementsGmail Find StatementsA

Find bank/card statement emails in Gmail (read-only). Returns message id/subject/ from/date/snippet for you to pick from. Omit query to use a default that targets common Indian bank/card statement senders in the last 3 months. Then call gmail_read_statement.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryNo
max_resultsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden. It clearly discloses the read-only nature, the returned fields (message id/subject/from/date/snippet), and the default query behavior, which is meaningful beyond the schema. It stops short of mentioning auth, rate limits, or query syntax, but for a read-only search this is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise, information-dense sentences. The first sentence states purpose and safety, the second describes output and selection intent, and the third explains the default behavior and next step. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with an output schema and a clearly named follow-up sibling, the description covers the essential workflow, default behavior, and return contents. The only notable omission is explicit explanation of the 'max_results' parameter and custom query usage, but the tool remains callable without those details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains the optional 'query' default behavior, but gives no detail on query syntax or semantics. 'max_results' is not mentioned at all, though its meaning is somewhat inferable from its name and default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Find') and resource ('bank/card statement emails in Gmail'), and clarifies it is read-only. It also names the follow-up tool (gmail_read_statement), which distinguishes this scanning step from the reading step.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context: the default query targets common Indian bank/card statement senders in the last 3 months, and the result should be picked from before calling gmail_read_statement. It does not explicitly state when not to use it, but the intended workflow is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gmail_read_statementGmail Read StatementA

Read one statement email's decoded text (read-only). Returns id/subject/from/date/text. Extract the transactions from text, then call import_statement to review + import them.

ParametersJSON Schema
NameRequiredDescriptionDefault
message_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the safety burden. It explicitly says 'read-only', which is the key behavioral disclosure, and mentions 'decoded text' and the returned fields. However, it does not disclose auth requirements, rate limits, or behavior when message_id is not a statement email.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no wasted words. The purpose is front-loaded, and the workflow instruction is placed directly after the return-value list. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists, so return structure is covered. The description adds the critical workflow step of calling import_statement. It is slightly incomplete in not mentioning the prerequisite step of obtaining message_id from gmail_find_statements or handling non-statement emails, but it is otherwise sufficient for a simple read operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It implies that message_id identifies the email to read, but it never explicitly defines the parameter or says where the ID comes from (e.g., a prior gmail_find_statements result). The parameter name is self-explanatory, but the origin is left to inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Read'), a specific resource ('one statement email's decoded text'), and explicitly marks the operation read-only. It also names the follow-up tool (import_statement), and the phrase 'one statement email' distinguishes it from the search-oriented sibling gmail_find_statements.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear workflow context: extract transactions from the returned text, then call import_statement to review and import them. It does not explicitly list when not to use it or name alternatives, but the workflow instruction strongly implies the intended usage chain.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_statementImport StatementA

Turn parsed statement rows into a reviewable import proposal (creates NOTHING).

transactions: list of {date: 'YYYY-MM-DD', merchant (or description), amount, category? (override), split_ref? (a saved default-split name)}. default_split_name: template applied to rows without their own split_ref; if omitted, rows default to 100%-personal (you pay + owe fully). dedup: if True, flags rows that match an existing expense in the local mirror (same day + same amount) and marks them to skip, so re-importing a statement doesn't double-add. (Run sync_all first for dedup to see your existing data.)

Returns proposals + summary. Review, then call confirm_import with the rows to create.

ParametersJSON Schema
NameRequiredDescriptionDefault
dedupNo
transactionsYes
default_split_nameNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and succeeds: it states that nothing is created, explains dedup's skip behavior to prevent double-adds, and defines default_split_name fallback semantics. This is strong behavioral disclosure for a non-mutating proposal tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well organized: one purpose line, then compact parameter explanations, then the follow-up instruction. Every sentence adds behavioral or workflow value; there is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a proposal-generation tool with an output schema, the description covers input shape, defaults, dedup side effects, required prerequisite, and next step. The only implicit content is the exact proposal structure, but that is covered by the available output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description manually documents all three parameters in detail: transaction item fields with date format and optional overrides, default_split_name behavior, and dedup matching rules. It fully compensates for the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line is specific: 'Turn parsed statement rows into a reviewable import proposal (creates NOTHING).' This names the exact operation and resource, and the parenthetical clearly distinguishes it from mutation tools like confirm_import or create_expense.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear workflow context: run sync_all first for dedup, then review the proposal, then call confirm_import. It also explains defaulting behavior for splits. It does not explicitly list alternative tools or say 'do not use create_expense,' but the staged workflow is unmistakable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_default_splitsList Default SplitsB

List all saved split templates.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It only states 'List all saved split templates' without mentioning ordering, pagination, or whether the returned list is sorted. There is no information about the output structure, though an output schema exists. The description adds minimal behavioral context beyond the basic operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with no wasted words. It is front-loaded with the action and resource, making it easy to parse. Perfectly concise for such a simple tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the zero-parameter interface and the presence of an output schema, the description is largely complete for a simple list operation. However, it does not mention whether the list is sorted or if there are any implicit filters. The absence of any caveats is acceptable for such a straightforward tool, but a note about the return format could enhance completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema fully covers parameter semantics (100% coverage). The description adds no parameter-specific information because there are none to describe. Per calibration, a baseline of 4 is appropriate when no parameters exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'list' and the resource 'saved split templates'. It is distinguishable from sibling tools like save_default_split and delete_default_split, which are mutation operations. However, it could be more explicit that this retrieves a collection of templates, not a single one.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no explicit guidance on when to use this tool versus alternatives. While it is implied that this is for retrieving existing split templates, there is no mention of when to prefer this over save_default_split or delete_default_split. The agent is left to infer the context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

remove_user_from_groupRemove User From GroupB

Remove a user from a group. User must have zero balance in the group.

ParametersJSON Schema
NameRequiredDescriptionDefault
user_idYes
group_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions the zero-balance constraint but does not state what happens if the user has a balance (e.g., error, failure), any authentication requirements, reversibility, or side effects. For a mutating operation, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely concise—two short sentences with zero fluff. The core action is front-loaded, and the constraint is stated immediately after. Every word earns its place, exemplifying efficient writing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having an output schema, the description is incomplete for a mutation tool of this nature. It lacks details on error handling (especially the zero-balance case), required permissions, and the exact effect of the operation. Given the low parameter explanation and absence of annotations, the agent is left without crucial operational context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate by explaining the parameters. It does not mention user_id or group_id at all, leaving the agent to infer their meaning from the names alone. No additional constraints, formats, or relationships are provided, failing to add any value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Remove a user from a group') and the resource involved, and it is easily distinguished from its sibling 'add_user_to_group' as the inverse operation. The verb 'remove' is unambiguous and directly tied to the resource, meeting the highest standard.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context (removal of a user) and a critical prerequisite ('User must have zero balance in the group'), which guides when it can be used. However, it does not mention alternatives or explicitly state when not to use this tool versus other group-related operations, so it stops short of full guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_categoryResolve CategoryB

Fuzzy-match an expense category by name (e.g. "food", "utilities"). Returns matches with id, name, and match_score. Searches subcategories too.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
thresholdNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavior. It does mention 'fuzzy-match' and 'searches subcategories,' which are useful behavioral traits, but it does not explicitly state that the operation is read-only or describe any side effects, error behavior, or limitations. For a tool that likely performs a safe lookup, this omission is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise—two sentences with no filler. It front-loads the core purpose and adds the subcategory detail efficiently. Every sentence contributes to understanding the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists, so return value details are not required. However, the tool has two parameters, and the threshold is unexplained. The description also lacks guidance on when to use this tool versus alternatives. For a moderate-complexity tool, this falls short of complete contextual guidance.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for explaining parameters. It indirectly explains the 'query' parameter by saying 'by name,' but it does not mention the 'threshold' parameter at all—its meaning, default behavior, or how it affects matching. This leaves a required parameter's semantics entirely undocumented beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: fuzzy-match an expense category by name, returns id, name, and match_score, and searches subcategories. It distinguishes from sibling tools like get_categories (which likely lists all categories) and resolve_friend/resolve_group by specifying the entity type and the fuzzy matching behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when you have a category name to resolve, but it does not explicitly state when to use this vs alternatives like get_categories, or when not to use it. It lacks guidance on prerequisites or exclusion criteria, leaving the agent to infer the appropriate context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_friendResolve FriendA

Fuzzy-match a friend by name. Returns matches with id, name, and match_score. Use this when you know a name but not the user_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
thresholdNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses that matching is fuzzy and that results include a match_score, which is helpful. However, it does not explain threshold behavior, result ordering, maximum matches, or what happens when no match is found.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences front-load the action and return type, then add the usage context. There is no filler or redundancy; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for a simple tool, especially with an output schema available. However, the unexplained threshold parameter and lack of behavioral details about fuzzy matching leave a meaningful gap for an agent trying to call it correctly in edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies the meaning of 'query' indirectly ('match a friend by name'), but it never explains the 'threshold' parameter, which is a nullable integer with a default of null and no schema description. This leaves a required parameter's semantics underspecified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Fuzzy-match a friend by name.' It also clarifies the return shape (id, name, match_score) and the triggering context ('know a name but not the user_id'), which distinguishes it from sibling tools like get_friend or get_friends.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states the intended use case: 'Use this when you know a name but not the user_id.' It gives clear context but does not mention exclusions or name alternative tools such as get_friend when an ID is known.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_groupResolve GroupA

Fuzzy-match a group by name. Returns matches with id, name, and match_score. Use this when you know a group name but not the group_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYes
thresholdNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that the operation is fuzzy matching and that results are scored, which is useful. However, it is silent on the threshold parameter behavior, handling of no matches, ordering, or determinism, leaving a notable gap in behavioral transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences: what it does, what it returns, and when to use it—all front-loaded with no filler. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is a simple lookup with an optional threshold, and the description covers the basic scenario (name → id). But the missing threshold semantics and lack of annotations leave the description incomplete for an agent that wants to control fuzzy-match strictness or anticipate edge cases. The output schema helps with return values but not invocation behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It implicitly explains 'query' as the group name, but says nothing about 'threshold', its default, or how it affects matching. With one of two parameters completely unexplained, the semantic contribution is inadequate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses a specific verb and resource ('fuzzy-match a group by name') and explicitly states the return shape ('matches with id, name, and match_score'). This clearly distinguishes it from exact lookups like get_group and sibling resolvers like resolve_friend or resolve_category, making its purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use this when you know a group name but not the group_id' is an explicit condition that tells an agent when to invoke the tool. It implicitly suggests alternatives for when the group_id is known, but it does not explicitly name those alternatives (e.g., get_group), so guidance is good but not fully explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

restore_expenseRestore ExpenseA

Restore a previously deleted expense. Use this to undo an accidental deletion.

ParametersJSON Schema
NameRequiredDescriptionDefault
expense_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden of behavioral disclosure. It conveys that the action is non-destructive and references a previously deleted expense, but it does not mention required permissions, idempotency, error behavior, or what happens if the expense_id is invalid or belongs to a non-deleted expense.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler. The core information—restoring a deleted expense and its intended use case—is front-loaded and immediately actionable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with an output schema, the description is nearly adequate, but it still lacks behavioral and error-context details that matter for safe invocation. The absence of annotations and undeclared permission or idempotency semantics leaves the description as only minimally viable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It vaguely implies that expense_id identifies a previously deleted expense, but it does not explicitly explain that this ID must refer to a deleted expense or clarify any constraints beyond the integer type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the action ('Restore') and the resource ('previously deleted expense'), making the tool's purpose unambiguous. It also establishes a clear contrast with the sibling delete_expense, so the agent can recognize this as the inverse operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Use this to undo an accidental deletion' provides a clear when-to-use condition. There are no competing restore tools among the siblings, so explicit exclusions or alternatives are not necessary, though they would round out the guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

save_default_splitSave Default SplitA

Save a reusable split template by name (e.g. "roomies-4way").

split: {"type": "equal"|"shares", "among": [user_id, ...], "shares"?: {user_id: weight}} Referenced from create_itemized_expense items via "split_ref": "". Stored locally in ~/.expensifyai/splits.json.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYes
splitYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does meaningful work: it states that data is persisted locally in ~/.expensifyai/splits.json and describes the reusable template structure. It does not mention overwrite behavior for an existing name, but the local storage path and template semantics provide substantial transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose, followed by a terse schema illustration and integration note. Every sentence adds useful information, with no filler or redundant restatement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter local persistence tool with an output schema, the description covers the tool's purpose, the split format, the storage location, and its relationship to create_itemized_expense. Nothing essential to invoking the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description fully compensates by explaining the name parameter with an example and detailing the split object's structure, including allowed type values, the among array, and the optional shares mapping. This is exactly the kind of parameter-level meaning the schema alone does not provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Save a reusable split template by name.' It clearly distinguishes this tool from the sibling list_default_splits and delete_default_split operations, which are the only other split-template tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly contextualizes usage by explaining that saved templates are 'Referenced from create_itemized_expense items via split_ref', which indicates when this tool is useful. It does not explicitly name alternatives or state when not to use it, but the integration context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_expensesSearch ExpensesA

Search the local mirror (run sync_all first). Full-text over description/details/ category, plus filters: amount range, user_id, group_id, category, date range. Returns matching expenses. Instant, offline, no API calls, no Pro paywall.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNo
user_idNo
categoryNo
group_idNo
max_amountNo
min_amountNo
dated_afterNo
dated_beforeNo
include_deletedNo
include_paymentsNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden. It clearly discloses that this is a local, offline, read-only operation with no API calls and no paywall, and it flags the sync dependency. It doesn't detail pagination or data-freshness edge cases, but the core behavior is well explained.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences with no fluff: purpose and dependency first, then search scope and filters, then return value and operational characteristics. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema and 11 optional, self-named parameters, the description covers the critical context: sync prerequisite, local/offline behavior, searchable fields, and filter dimensions. It could add a bit more about include_deleted/include_payments and limit semantics, but nothing blocks correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must add parameter meaning. It explains query as full-text over description/details/category and maps filters to amount range, user_id, group_id, category, and date range. A few optional params (limit, include_deleted, include_payments) are left to their self-explanatory names and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific operation ('Search the local mirror') and a clear resource (expenses), then details search scope and filters. The phrase 'local mirror' plus 'no API calls' distinguishes it from API-backed siblings like get_expenses.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a concrete prerequisite ('run sync_all first') and context for when the local search is appropriate ('Instant, offline, no API calls'). It does not explicitly name alternative tools or state when not to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_notesSearch NotesA

Fuzzy search over expense descriptions AND notes/details in the local mirror.

Use when exact search misses — catches typos, abbreviations (vizag/vskp/vtz), and amounts or words buried inside a bundled expense's note (e.g. a '5552' line inside a multi-item details field). Returns matches with a score (100 = exact substring) and whether it hit the description or the note. Run sync_all first.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
min_scoreNo
include_deletedNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses the local-mirror dependency, the sync prerequisite, and the return semantics ('score (100 = exact substring)' and 'whether it hit the description or the note'). It stops short of describing ordering or filtering behavior, but the output schema covers return structure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose, then gives use guidance, concrete examples, return behavior, and a prerequisite in three compact sentences. No sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a search tool with one required parameter and an output schema, the description covers the key invocation context: fuzzy behavior, the local mirror, sync requirement, and the two match targets. It could name search_expenses explicitly, but the 'exact search misses' phrasing makes the alternative clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must add meaning. It clarifies the query semantics thoroughly (typos, abbreviations, buried amounts/words) and explains the score scale, which gives meaning to min_score. However, limit and include_deleted are left to inference from their names/defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line names a specific verb and resource: 'Fuzzy search over expense descriptions AND notes/details in the local mirror.' It differentiates from siblings by explicitly framing this as the fallback when 'exact search misses' and by covering note details as well as descriptions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit trigger condition: 'Use when exact search misses — catches typos, abbreviations...' and gives examples of buried matches. It also states a prerequisite ('Run sync_all first'), so the agent knows the required precondition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

sync_allSync AllA

Sync Splitwise into a local SQLite mirror for instant offline search.

Delta sync: uses the API's updated_after cursor so only expenses that were added, edited, moved, or deleted since the last sync are fetched (first run, or full=True, pulls everything). Groups and friends are fully refreshed each run (small). Upserts by expense id, so re-running is safe. Returns counts.

ParametersJSON Schema
NameRequiredDescriptionDefault
fullNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite no annotations, the description discloses the delta-cursor mechanism, what data is refreshed each run, and what happens on re-run ('Upserts by expense id, so re-running is safe'). It also mentions 'Returns counts,' covering a practical side effect beyond the input schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three concise, information-dense sentences front-load the purpose, then detail sync behavior and safety. No filler; every sentence contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With one simple boolean parameter, no annotations, and an output schema present, the description covers the sync semantics, idempotence, and first-run behavior. Nothing an agent needs to decide or call correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, full, is undocumented in the schema (0% coverage), but the description explains it directly: 'first run, or full=True, pulls everything' and implies the default delta behavior. This fully compensates for the empty schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description opens with 'Sync Splitwise into a local SQLite mirror for instant offline search,' naming the verb, resource, and purpose. The delta-sync mechanism makes it distinct from sibling query tools like get_expenses or search_expenses, so an agent can differentiate it without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly frames the tool as the offline-mirror/sync option ('for instant offline search') and explains when a full pull happens ('first run, or full=True'), plus idempotence ('re-running is safe'). It does not explicitly name alternatives or give when-not-to-use guidance, but the context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unlock_statement_pdfUnlock Statement PdfA

Unlock + parse a password-protected statement PDF attached to a Gmail message.

Provide name + dob (+ optional card_last4) so passwords can be derived (common Indian bank formats), or pass an explicit password. Returns parsed transactions [{date, description, amount, credit}] plus counts. Then review via import_statement.

ParametersJSON Schema
NameRequiredDescriptionDefault
dobNo
nameNo
passwordNo
card_last4No
message_idYes
attachment_indexNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description transparently discloses the password-derivation logic ('common Indian bank formats') and the optional parameters that affect behavior. It also mentions the output structure, which is valuable for a tool with no annotations. It could be more explicit about side effects or failure modes, but it covers the key behavior well.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded with the core action, then details usage options and the next step. It is efficiently structured, though it could be slightly tightened without losing the useful context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 6 parameters and a complex password-derivation workflow, the description covers the necessary usage guidance and output structure. It doesn't explain every edge case, but it provides sufficient context for an agent to call it correctly, especially given the existing output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Despite 0% schema description coverage, the description explains how the parameters relate to password derivation (name + dob + optional card_last4) and offers the alternative of a direct password. It also references the return format, but does not detail the attachment_index parameter—however, the schema provides its default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action ('Unlock + parse'), the specific resource ('password-protected statement PDF attached to a Gmail message'), and its purpose (return parsed transactions). It differentiates from the sibling 'import_statement' by explicitly mentioning it in the final step, making the workflow clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides two clear usage paths: deriving passwords from personal info or providing an explicit password, and directs the user to 'review via import_statement' as a next step. However, it doesn't explicitly state when not to use this tool or list alternatives for other scenarios, such as using gmail_read_statement for non-protected PDFs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update_expenseUpdate ExpenseA

Update an existing expense. Only provided fields are changed. If any users are supplied, all shares for the expense are overwritten with the provided values.

ParametersJSON Schema
NameRequiredDescriptionDefault
costNo
dateNo
usersNo
detailsNo
group_idNo
expense_idYes
category_idNo
descriptionNo
currency_codeNo
repeat_intervalNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden. It goes beyond a generic 'update' by stating that unspecified fields are left untouched and that supplying users overwrites the entire shares mapping – a non-obvious destructive edge case. It doesn't cover permissions or side effects, but the most safety-relevant behaviors are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The primary verb phrase is front-loaded, and the critical edge case about users overwriting shares is isolated at the end for emphasis.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite the output schema and concise core semantics, the tool has 10 parameters and no annotations. The description leaves most parameter behavior, required identifiers beyond expense_id, and side effects undocumented, so an agent would still have to guess meaning for fields like repeat_interval or currency_code.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must add parameter-level meaning. It only explains the 'users' parameter's overwrite behavior; the other nine parameters (cost, date, details, group_id, category_id, description, currency_code, repeat_interval) are left to inference from their names and types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and object – 'Update an existing expense' – clearly identifying this as the mutation tool for an existing expense. The word 'existing' distinguishes it from create_expense, delete_expense, and get_expense without needing to inspect the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'Update an existing expense' implies when to use it: when an expense already exists and needs modification, not creation or deletion. However, it does not explicitly name alternatives, prerequisites, or exclusions such as when restore_expense or delete_expense would be more appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 42 tool updatesv0.2.0
    • First observedadd_user_to_group
    • First observedanalyze_spending
    • First observedattach_receipt
    • First observedcompare_group_members
    • First observedconfirm_import
    • First observedcreate_comment
    • First observedcreate_expense
    • First observedcreate_friend
    • First observedcreate_group
    • First observedcreate_itemized_expense
    • First observeddelete_comment
    • First observeddelete_default_split
    • First observeddelete_expense
    • First observeddelete_friend
    • First observeddelete_group
    • First observedget_categories
    • First observedget_comments
    • First observedget_currencies
    • First observedget_current_user
    • First observedget_expense
    • First observedget_expenses
    • First observedget_friend
    • First observedget_friends
    • First observedget_group
    • First observedget_groups
    • First observedget_notifications
    • First observedget_user
    • First observedgmail_find_statements
    • First observedgmail_read_statement
    • First observedimport_statement
    • First observedlist_default_splits
    • First observedremove_user_from_group
    • First observedresolve_category
    • First observedresolve_friend
    • First observedresolve_group
    • First observedrestore_expense
    • First observedsave_default_split
    • First observedsearch_expenses
    • First observedsearch_notes
    • First observedsync_all
    • First observedunlock_statement_pdf
    • First observedupdate_expense

TDQS

A3.6/5.0

Scored across 42 tools

Disambiguation4/5

Most tools have clear distinct purposes (expenses, groups, friends, splits, sync/search, import pipeline). Some potential confusion exists between get_expenses/search_expenses/search_notes and analyze_spending/compare_group_members, but descriptions clarify the differences.

Naming Consistency4/5

Tool names mostly follow a consistent verb_noun pattern (create_expense, get_expense, update_expense, delete_expense, restore_expense). Minor deviations: sync_all, analyze_spending, compare_group_members, and the gmail_* / unlock_statement_pdf tools break the pattern slightly but remain readable.

Tool Count3/5

42 tools is on the heavy side for a single server, but the breadth is justified by the domain (expense management, groups, friends, itemization, import pipeline, sync/search). Still, the count feels high and could overwhelm an agent.

Completeness5/5

The tool surface covers the full expense lifecycle (create, read, update, delete, restore), groups, friends, comments, categories, currencies, notifications, itemized expenses, default splits, receipt attachment, local sync/search, and a full statement import pipeline (Gmail → PDF unlock → import → confirm). No obvious dead ends.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI assistants to manage Splitwise expenses with atomic duplicate prevention, smart fuzzy matching, and support for flexible split ratios between two people.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables conversational control of Splitwise accounts through Claude AI, allowing users to add expenses, check group balances, record settlements, and manage payment splits using natural language commands. Supports multiple currencies and flexible splitting methods including equal, exact, and percentage-based divisions.
    1
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables natural language management of Splitwise expenses, groups, and friends via the Model Context Protocol, with dual authentication and fuzzy name resolution.
    13
    MIT