Tally
Enables Amazon Alexa+ to manage a shared household ledger by voice, including recording expenses, checking balances, settling debts, and managing household members.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@TallySam paid 132 dollars for dinner, split with Chris and Maya"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Tally — a shared household ledger you talk to
An MCP server for Alexa+ that tracks who paid for what in a household, and works out the fewest payments that settle everyone up.
"Alexa, Sam paid 132 dollars for dinner, split with Chris and Maya."
→ Recorded 132 dollars for dinner, paid by Sam, split 3 ways.
That's 44 dollars each.
"Alexa, who owes what?"
→ 2 payments settle Apartment 4B. The biggest: Maya owes Sam 44 dollars.Why voice, and why a shared device
Splitting expenses is not an arithmetic problem — a calculator solved that decades ago. It is a capture problem. Nobody unlocks a phone and opens an app at a restaurant table, so the expense goes unrecorded, and a week later the evening is settled from memory and somebody quietly eats the difference.
Voice removes the capture cost: you say it while the bill is still in your hand.
And Alexa is not a personal device — it sits in the kitchen and belongs to everyone in the apartment. A shared ledger on a shared device is something a per-phone app structurally cannot be: any flatmate can record a purchase or ask where things stand without installing anything.
Being shared is also why the payer gets named out loud. "Sam paid for dinner" works from the kitchen counter whoever is standing at it; "I paid" only means something on an account linked to one person, which is what the invite code establishes.
Related MCP server: CaseChaser for Alexa+
What it does
Tool | Spoken example |
| "Sam paid 132 for dinner, split with Chris and Maya" |
| "Who owes what?" |
| "Chris paid Sam back 44" |
| "No — cancel that" |
| "Add Dana to the apartment" |
| "Start a household called Apartment 4B" |
| "How much do I owe Chris?" |
| "What did we spend lately?" |
Three decisions that shaped the build
The spoken answer is the product; the screen is a bonus. Every tool returns
a sentence written to be heard once — it leads with what you should do, not
with a table nobody can hold in their head. client_supports_apps() decides
whether a screen exists at all: an Echo Show additionally gets the
MCP Apps
balance sheet, an Echo Dot loses nothing but the picture. Both paths are tested.
Never guess whose money it is. Speech recognition is lossy, so names get resolved exactly, then by unambiguous prefix. "An" with both Anna and Andrew in the apartment does not pick one — it asks, over MCP elicitation, and falls back to putting the question in the error text when the host has no back-channel. Misattributing an expense is worse than one extra question.
Money is integers, end to end. Every amount is a count of minor units, and floats never touch the arithmetic. A three-way split of 10.00 cannot be equal, so the odd cent is handed out on a rotation instead of always landing on the same person. Property-based tests assert the invariants that matter: balances always sum to zero, and the suggested transfers actually clear the ledger.
Architecture
tally/
money.py Minor-unit arithmetic; equal / weighted / exact splits
ledger.py Households, entries, balances, debt simplification
store.py SQLite persistence (WAL)
server.py MCP server: tools, Apps binding, elicitation
ui.py The ui:// app resource for screen devices
auth.py OAuth 2.1 authorization server (PKCE, rotation)
login.py Account linking — redeeming an invite codeThe ledger domain has no MCP dependency and no I/O, which is why its invariants can be hammered with generated input in milliseconds.
Settling up uses the standard greedy max-debtor/max-creditor match. Each
step zeroes at least one person, so it never needs more than n-1 transfers,
against the n(n-1)/2 of paying everybody back individually. (Finding the true
minimum is NP-hard; this is optimal unless a strict subset happens to balance
among itself.)
Running it
uv sync
uv run pytest # 243 tests
uv run python -m tally --stdio # for the MCP InspectorSeeing it work without an Echo
Alexa+ add-on publishing is limited to partners in the US, so the two scripts in
demo/ stand in for the device. Run all three in separate terminals:
uv run python -m tally --port 8000
uv run python demo/mock_host.py # http://127.0.0.1:8977
uv run python demo/conversation.py --speakconversation.py plays a scripted exchange against the running server: each
line is an utterance, the tool call Alexa+ would route it to, and the server's
own spoken answer — read aloud by the system voice, with per-call latency.
Nothing is stubbed; every reply came over Streamable HTTP.
mock_host.py renders the ui:// app at Echo Show 15 and Echo Show 5 sizes,
performing the same ui/initialize handshake a real host would and logging
every message the app sends back. It follows the live server, so the balance
sheet fills in as the conversation runs.
Speech and screen are rendered differently, on purpose
Alexa reads a tool's text aloud verbatim, so the same amount is written twice:
42 dollars 50 cents in the spoken answer, $42.50 in the structured payload
a screen renders. "42.50 USD" would be read out as "forty two point five zero
U S D".
Over HTTP, the transport Alexa+ uses:
uv run python -m tally --port 8000Connecting it to Alexa+
Alexa needs a public HTTPS origin and OAuth. Expose the local server and pass the public URL — that switch turns on discovery metadata, the authorization code flow with PKCE, and the account-linking page:
cloudflared tunnel --url http://localhost:8000
uv run python -m tally --port 8000 --public-url https://<your-tunnel>.trycloudflare.comAlexa+ does not register itself, so it needs a client configured up front.
Without TALLY_CLIENT_ID and TALLY_REDIRECT_URI (the values from the Alexa
developer console) account linking cannot start at all:
export TALLY_CLIENT_ID=... TALLY_REDIRECT_URI=https://layla.amazon.com/api/skill/link/...Then onboard the add-on:
alexa-ai configure
alexa-ai new mcp --name "Tally" --locale en-US \
--mcp-server-url https://<your-tunnel>.trycloudflare.com/mcp
alexa-ai deployWithout --public-url the server runs unauthenticated — fine for the Inspector
and the test suite, never for a real device.
Joining a household
A name is not a secret. Adding someone creates a placeholder member and a single-use invite code, which the person who set the household up passes on; linking a device redeems that code. Matching on the name instead would let anyone who guessed a flatmate's first name attach their own Alexa to somebody else's ledger.
Protocol conformance
MCP spec 2025-11-25+ over Streamable HTTP.
MCP Apps (
io.modelcontextprotocol/ui): aui://resource served astext/html;profile=mcp-app, bound to four tools via_meta.ui.resourceUri, with aui/initializehandshake and no network access at all in its CSP.OAuth 2.1:
/.well-known/oauth-authorization-server, authorization codePKCE (S256), dynamic client registration, refresh-token rotation, and RFC 8707 resource validation so a token minted for another server is refused.
Tool annotations: read-only and destructive hints, so a host knows what needs confirming before it runs.
Resources and prompts: the ledger is readable at
tally://householdas JSON — reading where things stand is not a side effect and should not need a tool call — and three prompts give a host somewhere to start.Change notifications: tools that move the ledger publish a resource update, so a listening host can refresh instead of going stale. This reaches
subscriptions/listenhosts (2026-07-28); at the 2025-11-25 version Alexa+ negotiates there is no push path at all — see FRICTION.md.Structured output: every ledger tool publishes an
outputSchemaand always returns structured content — what a screen changes is how much has to be said out loud, not what is sent.Latency: ~15 ms median per tool call over HTTP, against the 500 ms Alexa+ allows.
Testing
243 tests, in seven layers:
Property-based (Hypothesis) over the money and ledger invariants — conservation, fairness bounds, settlement correctness.
End-to-end over a real MCP client session — tool schemas, Apps negotiation, structured content, elicitation, error shape, persistence across sessions.
Hostile input — amounts that look numeric but are not, blank and enormous names, nonsense currencies, and shares that cannot be divided. Everything is refused with something a person can act on, and nothing reaches the database it cannot survive.
Concurrency — two dozen simultaneous writers: nothing is lost, the ledger still sums to zero, and no two entries claim the same place in the history.
Two people, one ledger — the claim the product rests on: a second flatmate joins, "I" means a different person for each of them, either can undo a misheard entry, and a stranger reaches nothing.
OAuth over real HTTP — the whole authorization code flow, plus the attacks it exists to stop: replayed codes and forged PKCE verifiers.
See FRICTION.md for what the Amazon and MCP tooling got right and where it cost time.
License
MIT — see LICENSE.
Available Tools
8 toolsadd_personAdd a personA
Add someone to the household so expenses can be split with them.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Their name | |
| also_called | No | Nicknames speech recognition might produce |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnlyHint=false, idempotentHint=false, and destructiveHint=false. The description adds 'Add' which is consistent with a non-read-only, non-idempotent write operation, but it does not provide additional behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that clearly communicates the tool's function and purpose without any unnecessary words or digressions.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description fully covers the tool's primary purpose and the context of household expense splitting. While it does not mention edge cases or output details, the simplicity of the operation means the description is sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides descriptions for both parameters (name: 'Their name', also_called: 'Nicknames speech recognition might produce'). The description does not add extra meaning beyond the schema, so baseline score of 3 is appropriate given 100% schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action (Add) and the resource (someone to the household), along with the purpose (so expenses can be split with them). This distinguishes it from sibling tools like record_expense or show_balances.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implicitly conveys when to use this tool: when you need to add a person to the household for expense splitting. It does not explicitly mention alternatives, but the use case is specific enough that an agent can infer appropriate usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recent_activityRecent activityBRead-onlyIdempotent
List what was recorded lately, most recent first.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many entries |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, establishing the safe read-only profile. The description adds the 'most recent first' ordering behavior, but does not disclose what kinds of records are included or how far back 'lately' reaches. There is no contradiction with the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence that front-loads the verb and includes the key ordering constraint. There is no filler, repetition of the title, or extraneous content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter read-only tool with annotations and an output schema, the description is mostly adequate. The main gap is that 'what was recorded lately' is ambiguous about scope—it does not specify whether this lists only expenses or all recorded household actions, so an agent may not know what results to expect.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single parameter limit is fully documented with a description, default, minimum, and maximum. The description adds no parameter-level detail, but it does not need to because the schema already carries the meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses the specific verb 'List' and names the resource ('what was recorded lately'), and adds the ordering detail 'most recent first.' This is enough to separate it from siblings like show_balances and record_expense, though 'what was recorded' is slightly broad and could include expenses, settlements, or other household actions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to use this tool versus alternatives such as show_balances, what_do_i_owe, or undo_last. The recency framing implies a use case, but the description leaves the selection decision entirely to the agent with no exclusions or alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_expenseRecord an expenseA
Record that someone paid for something on the household's behalf. This is the main tool: use it the moment the user mentions paying.
| Name | Required | Description | Default |
|---|---|---|---|
| amount | Yes | What was paid, e.g. '42.50' | |
| shares | No | Name to amount, when people owe different amounts. Must sum to the total. | |
| paid_by | No | Who paid. Omit when the speaker paid. | |
| description | Yes | What it was for, e.g. 'dinner' | |
| split_between | No | Who shares it. Omit to split across the whole household. | |
| idempotency_key | No | Repeat unchanged when resending; omit otherwise. |
Output Schema
| Name | Required | Description |
|---|---|---|
| settle | Yes | |
| summary | Yes | |
| balances | Yes | |
| household | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag non-read-only, non-idempotent, and non-destructive behavior, so the safety profile is covered. The description adds no side-effect detail such as altering household balances or interacting with undo_last, but it does clarify that the expense is on the household's behalf. Given the annotation coverage, this is adequate but thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and a direct usage trigger. Every word earns its place and there is no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a rich output schema and fully documented parameters, the description does not need to explain return values or every field. The trigger condition plus schema-level idempotency guidance is enough for correct invocation. The only minor gap is silence on the effect on balances, but that does not block usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema fully documents amount, paid_by, description, split_between, shares, and idempotency_key. The description only loosely mirrors these concepts ('someone paid for something') without adding format, constraint, or default details. The baseline of 3 is appropriate because the schema carries the parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear action and resource: record a payment made on the household's behalf. 'This is the main tool' signals it is the default expense-capture tool, which helps distinguish it from balance and activity siblings. It could be more explicit about not being for settlement or queries, but the core purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The instruction 'use it the moment the user mentions paying' is an explicit trigger condition. No exclusions or alternative routing are provided, but the sibling set makes the write-vs-read distinction easy to infer. This meets the clear-context level, just without formal when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
settle_upSettle upB
Record that one person paid another back.
| Name | Required | Description | Default |
|---|---|---|---|
| to | Yes | Who received the money | |
| amount | Yes | How much changed hands | |
| paid_by | No | Who paid. Omit when the speaker paid. | |
| idempotency_key | No | Repeat unchanged when resending; omit otherwise. |
Output Schema
| Name | Required | Description |
|---|---|---|
| settle | Yes | |
| summary | Yes | |
| balances | Yes | |
| household | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description only says 'record', which implies a write operation, but provides no details on side effects, idempotency, or whether balances are updated. Since annotations provide no hints (all false), the description carries the full burden but falls short.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that contains no unnecessary words. It is well-structured and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is minimal and does not explain what 'settle up' entails in a larger expense-tracking context, such as whether it updates balances or creates a transaction. This ambiguity could confuse an agent, especially with several sibling tools that overlap in purpose.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool description adds no additional meaning to the parameters; the schema already fully describes each parameter. With 100% schema coverage, the baseline is 3, and the description does not enhance it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: recording that one person paid another back. It is specific and distinct from sibling tools like record_expense, which handles shared expenses, making the tool's function unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as record_expense or what_do_i_owe. The description offers no context for choosing this tool over others.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
show_balancesWho owes whatARead-onlyIdempotent
Report every member's net position and the fewest payments that settle it.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| settle | Yes | |
| summary | Yes | |
| balances | Yes | |
| household | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already provide readOnlyHint=true and idempotentHint=true, so the description doesn't need to repeat that. It adds valuable context about the output scope (net positions and a minimal settlement plan), which goes beyond the annotations. No contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, well-structured sentence that front-loads the core action ('Report every member's net position') and then adds a concise qualifier about the settlement plan. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters, read-only/idempotent annotations, and an output schema present, the description sufficiently covers what the agent needs to invoke it correctly. The return format is likely covered by the output schema, so no additional explanation is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline of 4 applies. The description doesn't need to explain any parameters, and the schema confirms an empty argument list.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Report') and a clear resource ('every member's net position and the fewest payments that settle it'). It clearly distinguishes from siblings like 'what_do_i_owe', which likely targets an individual, whereas this covers all members.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'every member' implies a group-wide report, giving some context for when to use it. However, it does not explicitly state when to prefer this over alternatives like 'what_do_i_owe' for individual balances, nor does it mention any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
start_householdStart a householdA
Create a shared ledger and put the speaker in it as the first member.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | What to call it, e.g. 'Apartment 4B' or 'Tahoe trip' | |
| currency | No | ISO code, e.g. USD, EUR, GBP | USD |
| your_name | No | The speaker's own name | Me |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate readOnlyHint=false (mutation) and destructiveHint=false (not destructive). The description confirms it creates a new ledger and adds the speaker, adding the 'first member' detail beyond annotations. However, it does not disclose idempotency (what happens if called twice) or any error conditions, so it adds only modest context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence: 'Create a shared ledger and put the speaker in it as the first member.' It is front-loaded with the primary action and contains no fluff or redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema (so return values are covered elsewhere) and a clear action. The description fully explains what the tool does. It does not mention that this is typically the first step in using the household ledger or that calling it again might create another household, but these are inferable from the name and siblings. Overall, it is sufficiently complete for a creation tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: all three parameters (name, currency, your_name) have clear descriptions. The tool description does not add any parameter-specific meaning, but with full schema coverage, the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Create') and resource ('shared ledger'), and clarifies that the speaker becomes the first member. This clearly differentiates it from sibling tools like record_expense, show_balances, etc., which handle other aspects of the ledger.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for initial setup by describing creation of a new ledger and adding the speaker, but it does not explicitly state when to use this tool versus alternatives like add_person, nor does it mention any exclusions (e.g., if a household already exists). Guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
undo_lastUndo the last entryADestructive
Reverse the most recent expense or payment, for when something was misheard.
| Name | Required | Description | Default |
|---|---|---|---|
| idempotency_key | No | Repeat unchanged when resending; omit otherwise. |
Output Schema
| Name | Required | Description |
|---|---|---|
| settle | Yes | |
| summary | Yes | |
| balances | Yes | |
| household | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already flag destructiveHint=true, so the agent knows this is a destructive operation. The description reinforces that by saying it 'reverses' an entry. It does not explicitly state that only the most recent entry is affected (though that is implied by the name and description) or whether the undo itself can be undone, but this is a minor gap given the simplicity of the tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that states the action, the target, and a contextual trigger. There is no extraneous information, and the structure is clean and easy to parse. It fits the tool's simple nature perfectly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple undo operation, the description is complete. It tells what the tool does, when to use it, and (via annotations) that it is destructive. No output schema is provided, but that is acceptable given the tool's nature. An agent has enough context to decide when and how to invoke it without further clarification.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, idempotency_key, is fully described in the schema with a clear explanation: 'Repeat unchanged when resending; omit otherwise.' This is a standard idempotency key instruction and covers the parameter's purpose and usage. With 100% schema coverage and an accurate description, there is no ambiguity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Reverse the most recent expense or payment.' The verb 'Reverse' and resource 'most recent expense or payment' are specific and unambiguous. It also stands apart from sibling tools like record_expense or settle_up by being the only undo operation, so an agent can easily distinguish it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description includes an explicit usage scenario: 'for when something was misheard.' This tells the agent when to invoke the tool (e.g., after an incorrect voice entry). However, it could more broadly state 'when an incorrect entry was made' and does not mention any limitations or prerequisites (e.g., that an undoable entry must exist), so it leaves a little room for interpretation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
what_do_i_oweWhat do I oweARead-onlyIdempotent
Answer what the speaker owes one other person, or their overall position when no one is named. Use this for 'how much do I owe Chris'.
| Name | Required | Description | Default |
|---|---|---|---|
| person | No | Who to compare against. Omit for the overall position. |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, and the description adds useful directional semantics: the amount is what the speaker owes another person, not what is owed to them. It also clarifies the default behavior when no person is named, which is genuinely useful beyond the structured annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler, and the core behavior is front-loaded before the usage example. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple optional-parameter query tool, the description plus annotations and output schema cover selection, invocation, and behavioral expectations. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already documents the person parameter with 100% coverage, including that omitting it gives the overall position. The description reinforces this meaning but adds no new syntax, format, or edge-case detail beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('answer'), a precise resource (what the speaker owes), and scopes the behavior to one named person or the overall position when no one is named. The example 'how much do I owe Chris' makes its purpose unmistakable and distinguishes it from siblings like show_balances.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear usage context with an explicit example ('Use this for how much do I owe Chris') and explains the no-person case. It does not explicitly say when not to use it or name alternatives, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
add_person - First observed
recent_activity - First observed
record_expense - First observed
settle_up - First observed
show_balances - First observed
start_household - First observed
undo_last - First observed
what_do_i_owe
TDQS
Scored across 8 tools
Each tool has a clearly distinct purpose with no overlap. Record, show, settle, undo, start, add, query, and list actions are uniquely identifiable.
All names use snake_case and mostly follow verb_noun pattern. 'what_do_i_owe' is a phrase rather than a verb_noun, but it remains clear and consistent in style.
With 8 tools, the set is well-scoped for a household expense tracker, covering core actions without excess or deficiency.
Core lifecycle actions are covered, including recording, settling, undoing, household creation, member addition, and queries. Missing explicit edit/delete but undo_last partially compensates.
Maintenance
Related MCP Connectors
Track and split shared expenses across trips, events, and groups. Create groups, add expenses, and…
Split bills from your AI: read bills & balances, create equal splits, request settlements.
- ContamosOAuthxyz.contamos
Shared ledger for groups that share money: balances, expenses, transfers, budgets and reports.
Log expenses, receipts and mileage from chat: auto-categorise, split VAT, summarise, export, rebill.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceEnables users to track expenses by adding, listing, and summarizing them with category support through natural language.-
- AlicenseNot gradedqualityCmaintenanceGives Alexa+ voice access to track household insurance claims, refunds, repairs, and disputes, and to chase them via disclosed AI phone calls while keeping money and legal decisions with the human.MIT
- AlicenseNot gradedqualityCmaintenanceEnables voice-first shared object memory for Alexa+, letting users record where items were last reported, retrieve authorized locations, correct stale records, check shared items in and out, run guided Lost Mode searches, and manage privacy-aware access.MIT
- AlicenseNot gradedqualityCmaintenanceEnables Alexa+ to manage a shared household task list through natural conversation, including checking status, finding overdue or occasion-tagged tasks, adding items, marking them done, and reassigning owners.MIT