Skip to main content
Glama

Rail MCP

Rail is a budget, approval and receipt layer that agents call when they buy something on a human's behalf, dry-run by default, wrapping Stripe Link for settlement rather than issuing cards.

It is not a marketplace, storefront, or outreach tool. Rail owns the budget, the approval gate, and the receipt ledger.

Quickstart

Requires Node.js 20+. Run the published package with npx -y rail-mcp. Default mode is dry-run: no network and no charges.

Claude Desktop and any other host that takes an mcpServers block:

{
  "mcpServers": {
    "rail": {
      "command": "npx",
      "args": ["-y", "rail-mcp"]
    }
  }
}

Cursor: one-click install (install links). The config value is the base64 of {"command":"npx","args":["-y","rail-mcp"]}.

Install Rail in Cursor

cursor://anysphere.cursor-deeplink/mcp/install?name=rail&config=eyJjb21tYW5kIjoibnB4IiwiYXJncyI6WyIteSIsInJhaWwtbWNwIl19

Grok Bot: add a custom MCP with npx -y rail-mcp.

Do not put Stripe or Link secrets (STRIPE_SECRET_KEY, LINK_ACCESS_TOKEN, or live-spend flags) in this shared snippet. Real charges stay off unless those gates are set on purpose, outside the shared config. See Stripe Link seam.

Ledger files (budget.json, proposals.json, receipts.json) are written to ~/.rail in the user home directory, or to RAIL_DATA_DIR when that is set. They are local state, not part of the npm package. The default does not follow the process working directory, so a host that starts the server from / still writes under the home directory.

Related MCP server: agent-budget

Example

Set a budget, propose a purchase, have a human approve it, then read the receipt. Dry-run moves no money off the machine.

  1. set_budget → { "amount_usd": 50, "note": "week1" }

  2. propose_purchase → { "merchant": "Acme", "amount_usd": 12, "rationale": "need widgets" }

  3. get_budget still shows remaining 50.00 and open_proposals: 1. The proposal is pending until a human decides.

  4. decide_proposal → { "proposal_id": "prop_…", "decision": "approve" }

  5. get_receipts returns one settled_dry_run receipt with a settlement_ref. Remaining budget is 38.00.

Rejecting spends nothing. refund_receipt reverses a dry-run settlement and restores the budget.

Tools

Tool

Purpose

set_budget

Overwrite the local budget limit. Moves no money

get_budget

Read budget, spent, remaining, open proposals count

propose_purchase

Record a pending intent only. Moves no money. Needs a later human approval

list_proposals

Read proposals by pending (default) / approved / rejected / all

decide_proposal

Human confirmation step. approve settles (dry-run receipt + settlement_ref, remaining decremented). reject spends nothing. No double-decide

get_receipts

Read receipts, newest first

refund_receipt

Human confirmation step. Reverse a dry-run settlement and restore budget

Permission model

Proposing a purchase is cheap. Settling one is not. An agent may record intent without moving money. The human should confirm the call that settles a purchase or reverses a receipt.

Hosts should read title, description, and annotations from tools/list. Every hint is set explicitly. The spec defaults (readOnlyHint false, destructiveHint true, openWorldHint true) would otherwise treat every tool as a destructive open-world write. Rail's ledger is local. Default mode is dry-run: no network and no charge.

Tool

readOnlyHint

destructiveHint

idempotentHint

openWorldHint

Host should

set_budget

false

false

true

false

Allow. Overwrites the local limit only. The same arguments do not stack or spend

get_budget

true

false

true

false

Allow. Read-only

propose_purchase

false

false

false

false

Allow. Adds one pending intent and spends nothing. Each call creates a new proposal

list_proposals

true

false

true

false

Allow. Read-only

decide_proposal

false

true

true

false

Confirm with the human. approve settles and decrements remaining budget. A repeat does not settle again

get_receipts

true

false

true

false

Allow. Read-only

refund_receipt

false

true

true

false

Confirm with the human. Reverses a dry-run settlement and restores budget. A repeat does not refund again

destructiveHint is the confirmation signal, and it is true only for decide_proposal and refund_receipt. openWorldHint is false on every tool: dry-run never leaves the ledger, and live Link stays behind separate environment gates. Annotations are hints for the host. They do not themselves move money.

Mode defaults to RAIL_MODE=dry_run. Money is stored as integer cents; tools display USD with 2 decimals. IDs: prop_…, rcpt_…, dry-run Link refs lsrq_dry_… on settlement_ref.

Local checkout

tsx is a dev dependency for dogfood in this repo. npm run smoke runs the TypeScript sources directly and does not need a build. npm start compiles first, then runs dist/index.js (the same file the rail-mcp bin points at).

npm install
npm start          # tsc → dist/, then stdio MCP server
npm run dev        # tsx watch src/index.ts
npm run smoke      # dry-run ledger smoke test; throwaway dirs only (never ./data, ~/.rail, or RAIL_DATA_DIR)

npm run smoke forces RAIL_MODE=dry_run, never enables live charge gates, and writes the ledger only under temporary directories it creates and deletes. ./data, ~/.rail, and RAIL_DATA_DIR are left untouched even when RAIL_DATA_DIR is set in the environment. One check starts the server with its working directory at / and HOME pointed at a temp directory, then calls a tool. That call succeeds, and the ledger for that check is created under the temp home rather than /data or the real ~/.rail. An explicit RAIL_DATA_DIR still overrides the default in that launch.

Rail asks Link for a one-time credential. Rail still decides budget and policy and writes the receipt. This seam does not issue cards and does not render UI.

  • RAIL_MODE=dry_run (default, including when unset): createSpendRequest returns a fake Link spend-request id and status dry_run_pending_human. No Stripe or Link HTTP. Approve stores that id as settlement_ref on a settled_dry_run receipt and decrements remaining budget.

  • RAIL_MODE=live without gates: throws live mode not enabled — set keys and get explicit human approval unless both STRIPE_SECRET_KEY and RAIL_LIVE=1 are set. The proposal stays pending and no receipt is written.

  • Live with those two gates: the POST https://api.link.com/spend_requests body is scaffolded only (see src/stripeLink.ts). Nothing is sent unless RAIL_ALLOW_LIVE_CHARGE=1.

  • RAIL_ALLOW_LIVE_CHARGE=1: the only branch allowed to call Link, and only if LINK_ACCESS_TOKEN is also set. STRIPE_SECRET_KEY is a presence gate and is never sent. There is no Stripe Issuing card create. Docs: Link CLI, link-cli, Issuing for agents.

  • Live spend stays off until those env vars are set on purpose. A posted Link request is stored as link_pending_human and does not decrement the dry-run budget. CI and npm run smoke stay on dry-run and do not charge.

Env

Role

RAIL_MODE

dry_run (default) or live

RAIL_DATA_DIR

Ledger directory. Default is ~/.rail in the user home directory. An explicit value overrides the default

STRIPE_SECRET_KEY

Required for live. Not used to issue cards or as a Link bearer token

RAIL_LIVE=1

Explicit live approval alongside the secret key

RAIL_ALLOW_LIVE_CHARGE=1

Additional gate before any Link HTTP

LINK_ACCESS_TOKEN

Link OAuth token. Required before the scaffolded POST actually runs

Explicit non-goals

  • No marketplace UI

  • No outreach / messaging

  • No card issuing

  • No real charges in dry-run, CI, or smoke

License

MIT. Copyright 2026 Thomas Stockham.

Available Tools

7 tools
decide_proposalDecide proposalA
DestructiveIdempotent

Approve or reject one pending proposal. This is the step a human should confirm. decision=approve settles the purchase: in the default dry-run mode it writes a settled_dry_run receipt, stamps a fake settlement ref (lsrq_dry_…), and decrements the remaining budget, without calling the network. decision=reject spends nothing and writes no receipt. The proposal can be decided only once; repeating the call does not settle again. Live charging stays off unless separate live gates are set on purpose.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNoOptional decision note
decisionYesapprove settles; reject spends nothing
proposal_idYesProposal id (prop_…)

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well past the annotations: it discloses dry-run behavior (settled_dry_run receipt, fake lsrq_dry_… settlement ref), budget decrement, that reject spends nothing and writes no receipt, and that live charging is gated. These details are consistent with destructiveHint=true, idempotentHint=true, and readOnlyHint=false, and they explain exactly what the mutation destroys or creates.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the action and the human-confirmation framing, then descending into dry-run mechanics and idempotency. No filler and every clause carries information an agent needs before calling.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must carry the effect story — and it does, describing receipt creation, settlement ref format, budget impact, and repeat-call behavior. For a destructive, non-open-world mutation with full annotation coverage, nothing an agent needs is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3, but the description adds real meaning beyond the enum text by spelling out the side effects of each decision value (approve settles, rejects spends nothing, writes no receipt). It does not explain note or proposal_id handling, so it does not fully compensate into a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb pair (approve/reject) on a specific resource (one pending proposal) and scopes it to exactly one item. This clearly separates it from propose_purchase (creates) and list_proposals (reads), so an agent can route without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"This is the step a human should confirm" gives a clear use condition, and the approve/reject branch guidance tells the agent which input to pick for which outcome. It does not explicitly name sibling alternatives (e.g., refund_receipt for unwinding, deliver_receipt style flows), so it stops short of full when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_budgetGet budgetA
Read-onlyIdempotent

Read the active budget, amount spent, amount remaining, and the count of pending proposals. Does not change the ledger, spend money, or contact Stripe or Link.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, non-destructive, and closed-world. The description adds value beyond that by ruling out external side effects with Stripe or Link and by describing what is returned. It does not cover pagination or freshness, but the extra context is meaningful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tight sentence front-loads what is read and follows with the explicit exclusions. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-param, read-only tool with no output schema, the description is nearly sufficient: it enumerates the returned fields and rules out side effects. It could note freshness or failure modes, but an agent has enough to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. The description compensates slightly by listing the values returned in place of the absent output schema, though there is no parameter semantics to clarify.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (read) and resource (budget) and enumerates the returned values (amount spent, remaining, pending proposals count), which lets an agent distinguish it from the mutation siblings. It stops short of naming an alternative like set_budget explicitly, but the read scope is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides negative guidance ('does not change the ledger, spend money, or contact Stripe or Link') that implicitly separates it from set_budget, propose_purchase, and refund_receipt, but never states when to call it or names a sibling to prefer instead. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_receiptsGet receiptsA
Read-onlyIdempotent

Read receipts, newest first, including dry-run settlements and refunds. Does not refund, spend money, or contact Stripe or Link.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax receipts to return

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint=false, so the safety profile is covered. The description adds genuine context beyond that: result ordering, inclusion of dry-run settlements and refunds, and the fact that no external calls (Stripe/Link) are made, reinforcing the closed-world hint. It omits pagination/limit behavior, which keeps it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and scope, then the disambiguating negatives. No redundant restatement of the title or name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read tool with full annotation coverage and no output schema, the description supplies ample scope (ordering, included record types). The only omission is return-size/pagination expectations, a minor gap rather than a blocking one.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Only one parameter ('limit') and schema coverage is 100%, so the schema already documents it. The description adds nothing about the limit parameter or default result size; per the rubric, baseline 3 applies when the schema fully carries parameter meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Read') plus resource ('receipts'), with ordering ('newest first') and content scope ('including dry-run settlements and refunds'). The closing negative clause ('Does not refund, spend money, or contact Stripe or Link') cleanly separates it from the refund_receipt sibling without the agent opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The explicit 'does not refund, spend money' carve-out tells the agent when to route elsewhere (refund_receipt / purchase flows), which is real routing guidance. It stops short of naming the alternative tool or stating a positive use trigger like 'use to audit past transactions', so it is clear context rather than a full when/when-not map.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_proposalsList proposalsA
Read-onlyIdempotent

Read purchase proposals filtered by status: pending (default), approved, rejected, or all. Does not approve, reject, spend money, or contact Stripe or Link.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusNoFilter: pending (default), approved, rejected, or all

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false. The description reinforces and extends this by stating the tool will not approve, reject, spend money, or contact Stripe/Link, giving concrete behavioral guardrails beyond the raw hints. It does not cover result volume or pagination, so it is good but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero filler; the primary action and filter options are front-loaded and the exclusion clause follows. Every clause carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single optional enum parameter with full schema coverage and a mature annotation set, the description gives enough to invoke correctly, including the null-argument default behavior. No output schema exists but return format need not be described; only result volume/pagination is unaddressed, which is minor here.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single enum parameter is fully documented in the schema, including the 'pending' default. The description restates the same enum values without adding syntax, format, or edge-case meaning, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read purchase proposals') plus the filtering dimension ('by status'). The negative clause ('Does not approve, reject, spend money, or contact Stripe or Link') implicitly separates it from decide_proposal and propose_purchase, so an agent can distinguish it from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Clearly establishes the read-vs-act context by enumerating statuses and explicitly ruling out approval, rejection, spending, and external contact, which routes the agent elsewhere for those actions. It stops short of naming the sibling tool (decide_proposal) that handles approval/rejection, so it is clear context without an explicit alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

propose_purchasePropose purchaseA

Record a pending purchase intent only. Moves no money, creates no receipt, approves no card, and does not contact Stripe or Link. Needs a later human approval via decide_proposal before anything settles. When a budget exists and the amount is above what remains, the proposal is stored with over_budget set and still spends nothing.

ParametersJSON Schema
NameRequiredDescriptionDefault
skuNoSKU or product identifier
urlNoProduct or checkout URL
merchantYesMerchant or seller name
rationaleNoWhy the agent wants this
amount_usdYesPurchase amount in USD

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With annotation coverage limited to generic hints (readOnlyHint=false, destructiveHint=false, idempotentHint=false), the description carries the real behavioral burden and does: it enumerates the negative guarantees (no money moved, no receipt, no card approval, no Stripe/Link contact), states the human-approval dependency, and discloses the over_budget edge-case behavior. This is exactly the side-effect disclosure a mutation tool needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, no filler, and the strongest constraint (intent-only, no money) is front-loaded ahead of the approval dependency and the edge case. The final over_budget sentence is edge-case detail that is useful but slightly beyond the minimum, keeping it short of a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by describing the observable outcome (proposal stored, over_budget set when the amount exceeds remaining budget). Combined with the safety profile and the decide_proposal hand-off, an agent has everything needed to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all five parameters are already documented in the schema and the description adds no format, range, or validation detail beyond it. Baseline 3 is appropriate when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Record a pending purchase intent') and immediately scopes it as intent-only, sharply distinguishing it from decide_proposal, which performs the actual approval. An agent can tell it apart from its siblings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It is explicit that this records intent only and that a later human approval via decide_proposal is required before anything settles, which routes the agent to the correct next tool. It does not state an explicit when-not condition (e.g., when to skip proposing and use get_budget first), but the 'only' scoping is a clear usage boundary.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

refund_receiptRefund receiptA
DestructiveIdempotent

Reverse a settled dry-run receipt and restore that amount to the remaining budget. This is the step a human should confirm. Does not contact Stripe or Link. A receipt can be refunded only once; repeating the call does not refund again. Refuses receipts that are not settled_dry_run.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonNoRefund reason
receipt_idYesReceipt id (rcpt_…)

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations: it discloses that no external call is made ('Does not contact Stripe or Link'), that the budget is restored, that the operation is once-only (reinforcing idempotentHint=true), and the exact refusal condition. Annotations cover the mutation/idempotency profile, and the description adds the missing operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, zero filler: the action is front-loaded, followed by the human-confirmation requirement and the two key behavioral constraints. Every sentence carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive mutation with no output schema, the description covers the essentials: what it reverses, the budget side effect, the external-call behavior, the once-only semantics, and the refusal precondition. Nothing needed to call it safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and there are only 2 parameters, so the schema already documents receipt_id and reason. The description adds no format, constraint, or meaning for either parameter beyond what is already encoded, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Reverse') and resource ('settled dry-run receipt') plus the resulting side effect ('restore that amount to the remaining budget'). This is clearly distinguishable from the sibling budget/proposal tools without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly flags that this is 'the step a human should confirm,' which tells the agent when to pause for approval, and states the precondition ('Refuses receipts that are not settled_dry_run'). It does not name an alternative tool, but none of the siblings overlaps, so this is a clear context rather than a full routing rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_budgetSet budgetA
Idempotent

Overwrite the local spend budget. Changes the ledger limit only. Does not spend money, approve a card, settle a purchase, or contact Stripe or Link. Passing the same amount again does not add to the budget.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNoOptional note for this budget
currencyNoCurrency code; default USD
amount_usdYesBudget amount in USD

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare idempotentHint=true and openWorldHint=false, so the description's job is to add nuance. It does so usefully: 'Passing the same amount again does not add to the budget' clarifies that idempotency here means overwrite rather than increment, and 'Changes the ledger limit only' scopes the blast radius. It omits any permission or auth requirements, keeping it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and effect, then the exclusions and the idempotency caveat. Every sentence carries distinct information with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter mutation with no output schema and annotations covering the safety profile, the description covers what the tool changes, what it does not change, and repeat-call behavior. It lacks any statement of required permissions or the resulting state after the overwrite, which keeps it from a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so amount_usd, currency, and note are already documented in the schema; the baseline of 3 applies. The description adds no format, range, or unit detail beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Overwrite the local spend budget') plus the precise effect ('Changes the ledger limit only'). It is clearly distinguishable from siblings like propose_purchase and refund_receipt, which the description explicitly disclaims.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The negation list ('Does not spend money, approve a card, settle a purchase, or contact Stripe or Link') effectively tells the agent when this tool is not the right one, which is strong routing guidance. It stops short of naming the sibling tool an agent should use instead, so it is not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv0.1.1
    • First observeddecide_proposal
    • First observedget_budget
    • First observedget_receipts
    • First observedlist_proposals
    • First observedpropose_purchase
    • First observedrefund_receipt
    • First observedset_budget

TDQS

A4.3/5.0

Scored across 7 tools

Disambiguation5/5

Each tool maps to a distinct resource+action: budgets (set/get), proposals (propose/list/decide), and receipts (get/refund). Descriptions explicitly clarify boundaries, e.g. set_budget vs propose_purchase vs decide_proposal, leaving no real overlap.

Naming Consistency5/5

Every tool follows a clean verb_noun snake_case pattern (set_budget, get_budget, propose_purchase, list_proposals, decide_proposal, get_receipts, refund_receipt). No mixing of conventions.

Tool Count5/5

Seven tools is well-scoped for a spend-governance ledger, each covering a necessary stage of the budget/proposal/receipt lifecycle with no filler.

Completeness4/5

Covers the full lifecycle: budget set/read, proposal create/list/decide, receipt read/refund. Minor gaps like filtering or clearing stale proposals exist but core workflows have no dead ends.

Related MCP Connectors

Related MCP Servers