rail-mcp
Integrates with Stripe Link as the settlement seam for approved purchases. It can request Link spend requests/one-time credentials, record dry-run settlement references, and supports live-mode Link HTTP calls behind explicit environment gates, while Rail itself manages budget, approvals, and receipts.
Rail MCP
Rail is a budget, approval and receipt layer that agents call when they buy something on a human's behalf, dry-run by default, wrapping Stripe Link for settlement rather than issuing cards.
It is not a marketplace, storefront, or outreach tool. Rail owns the budget, the approval gate, and the receipt ledger.
Quickstart
Requires Node.js 20+. Run the published package with npx -y rail-mcp. Default mode is dry-run: no network and no charges.
Claude Desktop and any other host that takes an mcpServers block:
{
"mcpServers": {
"rail": {
"command": "npx",
"args": ["-y", "rail-mcp"]
}
}
}Cursor: one-click install (install links). The config value is the base64 of {"command":"npx","args":["-y","rail-mcp"]}.
cursor://anysphere.cursor-deeplink/mcp/install?name=rail&config=eyJjb21tYW5kIjoibnB4IiwiYXJncyI6WyIteSIsInJhaWwtbWNwIl19Grok Bot: add a custom MCP with npx -y rail-mcp.
Do not put Stripe or Link secrets (STRIPE_SECRET_KEY, LINK_ACCESS_TOKEN, or live-spend flags) in this shared snippet. Real charges stay off unless those gates are set on purpose, outside the shared config. See Stripe Link seam.
Ledger files (budget.json, proposals.json, receipts.json) are written to ~/.rail in the user home directory, or to RAIL_DATA_DIR when that is set. They are local state, not part of the npm package. The default does not follow the process working directory, so a host that starts the server from / still writes under the home directory.
Related MCP server: agent-budget
Example
Set a budget, propose a purchase, have a human approve it, then read the receipt. Dry-run moves no money off the machine.
set_budget→{ "amount_usd": 50, "note": "week1" }propose_purchase→{ "merchant": "Acme", "amount_usd": 12, "rationale": "need widgets" }get_budgetstill shows remaining50.00andopen_proposals: 1. The proposal is pending until a human decides.decide_proposal→{ "proposal_id": "prop_…", "decision": "approve" }get_receiptsreturns onesettled_dry_runreceipt with asettlement_ref. Remaining budget is38.00.
Rejecting spends nothing. refund_receipt reverses a dry-run settlement and restores the budget.
Tools
Tool | Purpose |
| Overwrite the local budget limit. Moves no money |
| Read budget, spent, remaining, open proposals count |
| Record a pending intent only. Moves no money. Needs a later human approval |
| Read proposals by |
| Human confirmation step. |
| Read receipts, newest first |
| Human confirmation step. Reverse a dry-run settlement and restore budget |
Permission model
Proposing a purchase is cheap. Settling one is not. An agent may record intent without moving money. The human should confirm the call that settles a purchase or reverses a receipt.
Hosts should read title, description, and annotations from tools/list. Every hint is set explicitly. The spec defaults (readOnlyHint false, destructiveHint true, openWorldHint true) would otherwise treat every tool as a destructive open-world write. Rail's ledger is local. Default mode is dry-run: no network and no charge.
Tool | readOnlyHint | destructiveHint | idempotentHint | openWorldHint | Host should |
| false | false | true | false | Allow. Overwrites the local limit only. The same arguments do not stack or spend |
| true | false | true | false | Allow. Read-only |
| false | false | false | false | Allow. Adds one pending intent and spends nothing. Each call creates a new proposal |
| true | false | true | false | Allow. Read-only |
| false | true | true | false | Confirm with the human. |
| true | false | true | false | Allow. Read-only |
| false | true | true | false | Confirm with the human. Reverses a dry-run settlement and restores budget. A repeat does not refund again |
destructiveHint is the confirmation signal, and it is true only for decide_proposal and refund_receipt. openWorldHint is false on every tool: dry-run never leaves the ledger, and live Link stays behind separate environment gates. Annotations are hints for the host. They do not themselves move money.
Mode defaults to RAIL_MODE=dry_run. Money is stored as integer cents; tools display USD with 2 decimals. IDs: prop_…, rcpt_…, dry-run Link refs lsrq_dry_… on settlement_ref.
Local checkout
tsx is a dev dependency for dogfood in this repo. npm run smoke runs the TypeScript sources directly and does not need a build. npm start compiles first, then runs dist/index.js (the same file the rail-mcp bin points at).
npm install
npm start # tsc → dist/, then stdio MCP server
npm run dev # tsx watch src/index.ts
npm run smoke # dry-run ledger smoke test; throwaway dirs only (never ./data, ~/.rail, or RAIL_DATA_DIR)npm run smoke forces RAIL_MODE=dry_run, never enables live charge gates, and writes the ledger only under temporary directories it creates and deletes. ./data, ~/.rail, and RAIL_DATA_DIR are left untouched even when RAIL_DATA_DIR is set in the environment. One check starts the server with its working directory at / and HOME pointed at a temp directory, then calls a tool. That call succeeds, and the ledger for that check is created under the temp home rather than /data or the real ~/.rail. An explicit RAIL_DATA_DIR still overrides the default in that launch.
Stripe Link seam
Rail asks Link for a one-time credential. Rail still decides budget and policy and writes the receipt. This seam does not issue cards and does not render UI.
RAIL_MODE=dry_run(default, including when unset):createSpendRequestreturns a fake Link spend-request id and statusdry_run_pending_human. No Stripe or Link HTTP. Approve stores that id assettlement_refon asettled_dry_runreceipt and decrements remaining budget.RAIL_MODE=livewithout gates: throwslive mode not enabled — set keys and get explicit human approvalunless bothSTRIPE_SECRET_KEYandRAIL_LIVE=1are set. The proposal stays pending and no receipt is written.Live with those two gates: the
POST https://api.link.com/spend_requestsbody is scaffolded only (seesrc/stripeLink.ts). Nothing is sent unlessRAIL_ALLOW_LIVE_CHARGE=1.RAIL_ALLOW_LIVE_CHARGE=1: the only branch allowed to call Link, and only ifLINK_ACCESS_TOKENis also set.STRIPE_SECRET_KEYis a presence gate and is never sent. There is no Stripe Issuing card create. Docs: Link CLI, link-cli, Issuing for agents.Live spend stays off until those env vars are set on purpose. A posted Link request is stored as
link_pending_humanand does not decrement the dry-run budget. CI andnpm run smokestay on dry-run and do not charge.
Env | Role |
|
|
| Ledger directory. Default is |
| Required for live. Not used to issue cards or as a Link bearer token |
| Explicit live approval alongside the secret key |
| Additional gate before any Link HTTP |
| Link OAuth token. Required before the scaffolded POST actually runs |
Explicit non-goals
No marketplace UI
No outreach / messaging
No card issuing
No real charges in dry-run, CI, or smoke
License
MIT. Copyright 2026 Thomas Stockham.
Available Tools
7 toolsdecide_proposalDecide proposalADestructiveIdempotent
Approve or reject one pending proposal. This is the step a human should confirm. decision=approve settles the purchase: in the default dry-run mode it writes a settled_dry_run receipt, stamps a fake settlement ref (lsrq_dry_…), and decrements the remaining budget, without calling the network. decision=reject spends nothing and writes no receipt. The proposal can be decided only once; repeating the call does not settle again. Live charging stays off unless separate live gates are set on purpose.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | Optional decision note | |
| decision | Yes | approve settles; reject spends nothing | |
| proposal_id | Yes | Proposal id (prop_…) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well past the annotations: it discloses dry-run behavior (settled_dry_run receipt, fake lsrq_dry_… settlement ref), budget decrement, that reject spends nothing and writes no receipt, and that live charging is gated. These details are consistent with destructiveHint=true, idempotentHint=true, and readOnlyHint=false, and they explain exactly what the mutation destroys or creates.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the action and the human-confirmation framing, then descending into dry-run mechanics and idempotency. No filler and every clause carries information an agent needs before calling.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must carry the effect story — and it does, describing receipt creation, settlement ref format, budget impact, and repeat-call behavior. For a destructive, non-open-world mutation with full annotation coverage, nothing an agent needs is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds real meaning beyond the enum text by spelling out the side effects of each decision value (approve settles, rejects spends nothing, writes no receipt). It does not explain note or proposal_id handling, so it does not fully compensate into a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (approve/reject) on a specific resource (one pending proposal) and scopes it to exactly one item. This clearly separates it from propose_purchase (creates) and list_proposals (reads), so an agent can route without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"This is the step a human should confirm" gives a clear use condition, and the approve/reject branch guidance tells the agent which input to pick for which outcome. It does not explicitly name sibling alternatives (e.g., refund_receipt for unwinding, deliver_receipt style flows), so it stops short of full when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_budgetGet budgetARead-onlyIdempotent
Read the active budget, amount spent, amount remaining, and the count of pending proposals. Does not change the ledger, spend money, or contact Stripe or Link.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, non-destructive, and closed-world. The description adds value beyond that by ruling out external side effects with Stripe or Link and by describing what is returned. It does not cover pagination or freshness, but the extra context is meaningful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence front-loads what is read and follows with the explicit exclusions. Every clause earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-param, read-only tool with no output schema, the description is nearly sufficient: it enumerates the returned fields and rules out side effects. It could note freshness or failure modes, but an agent has enough to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. The description compensates slightly by listing the values returned in place of the absent output schema, though there is no parameter semantics to clarify.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (read) and resource (budget) and enumerates the returned values (amount spent, remaining, pending proposals count), which lets an agent distinguish it from the mutation siblings. It stops short of naming an alternative like set_budget explicitly, but the read scope is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides negative guidance ('does not change the ledger, spend money, or contact Stripe or Link') that implicitly separates it from set_budget, propose_purchase, and refund_receipt, but never states when to call it or names a sibling to prefer instead. Usage is inferable rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_receiptsGet receiptsARead-onlyIdempotent
Read receipts, newest first, including dry-run settlements and refunds. Does not refund, spend money, or contact Stripe or Link.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max receipts to return |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and openWorldHint=false, so the safety profile is covered. The description adds genuine context beyond that: result ordering, inclusion of dry-run settlements and refunds, and the fact that no external calls (Stripe/Link) are made, reinforcing the closed-world hint. It omits pagination/limit behavior, which keeps it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action and scope, then the disambiguating negatives. No redundant restatement of the title or name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read tool with full annotation coverage and no output schema, the description supplies ample scope (ordering, included record types). The only omission is return-size/pagination expectations, a minor gap rather than a blocking one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter ('limit') and schema coverage is 100%, so the schema already documents it. The description adds nothing about the limit parameter or default result size; per the rubric, baseline 3 applies when the schema fully carries parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Read') plus resource ('receipts'), with ordering ('newest first') and content scope ('including dry-run settlements and refunds'). The closing negative clause ('Does not refund, spend money, or contact Stripe or Link') cleanly separates it from the refund_receipt sibling without the agent opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The explicit 'does not refund, spend money' carve-out tells the agent when to route elsewhere (refund_receipt / purchase flows), which is real routing guidance. It stops short of naming the alternative tool or stating a positive use trigger like 'use to audit past transactions', so it is clear context rather than a full when/when-not map.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_proposalsList proposalsARead-onlyIdempotent
Read purchase proposals filtered by status: pending (default), approved, rejected, or all. Does not approve, reject, spend money, or contact Stripe or Link.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Filter: pending (default), approved, rejected, or all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false. The description reinforces and extends this by stating the tool will not approve, reject, spend money, or contact Stripe/Link, giving concrete behavioral guardrails beyond the raw hints. It does not cover result volume or pagination, so it is good but not exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler; the primary action and filter options are front-loaded and the exclusion clause follows. Every clause carries distinct information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single optional enum parameter with full schema coverage and a mature annotation set, the description gives enough to invoke correctly, including the null-argument default behavior. No output schema exists but return format need not be described; only result volume/pagination is unaddressed, which is minor here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single enum parameter is fully documented in the schema, including the 'pending' default. The description restates the same enum values without adding syntax, format, or edge-case meaning, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read purchase proposals') plus the filtering dimension ('by status'). The negative clause ('Does not approve, reject, spend money, or contact Stripe or Link') implicitly separates it from decide_proposal and propose_purchase, so an agent can distinguish it from siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clearly establishes the read-vs-act context by enumerating statuses and explicitly ruling out approval, rejection, spending, and external contact, which routes the agent elsewhere for those actions. It stops short of naming the sibling tool (decide_proposal) that handles approval/rejection, so it is clear context without an explicit alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
propose_purchasePropose purchaseA
Record a pending purchase intent only. Moves no money, creates no receipt, approves no card, and does not contact Stripe or Link. Needs a later human approval via decide_proposal before anything settles. When a budget exists and the amount is above what remains, the proposal is stored with over_budget set and still spends nothing.
| Name | Required | Description | Default |
|---|---|---|---|
| sku | No | SKU or product identifier | |
| url | No | Product or checkout URL | |
| merchant | Yes | Merchant or seller name | |
| rationale | No | Why the agent wants this | |
| amount_usd | Yes | Purchase amount in USD |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With annotation coverage limited to generic hints (readOnlyHint=false, destructiveHint=false, idempotentHint=false), the description carries the real behavioral burden and does: it enumerates the negative guarantees (no money moved, no receipt, no card approval, no Stripe/Link contact), states the human-approval dependency, and discloses the over_budget edge-case behavior. This is exactly the side-effect disclosure a mutation tool needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, no filler, and the strongest constraint (intent-only, no money) is front-loaded ahead of the approval dependency and the edge case. The final over_budget sentence is edge-case detail that is useful but slightly beyond the minimum, keeping it short of a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by describing the observable outcome (proposal stored, over_budget set when the amount exceeds remaining budget). Combined with the safety profile and the decide_proposal hand-off, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are already documented in the schema and the description adds no format, range, or validation detail beyond it. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Record a pending purchase intent') and immediately scopes it as intent-only, sharply distinguishing it from decide_proposal, which performs the actual approval. An agent can tell it apart from its siblings without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It is explicit that this records intent only and that a later human approval via decide_proposal is required before anything settles, which routes the agent to the correct next tool. It does not state an explicit when-not condition (e.g., when to skip proposing and use get_budget first), but the 'only' scoping is a clear usage boundary.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
refund_receiptRefund receiptADestructiveIdempotent
Reverse a settled dry-run receipt and restore that amount to the remaining budget. This is the step a human should confirm. Does not contact Stripe or Link. A receipt can be refunded only once; repeating the call does not refund again. Refuses receipts that are not settled_dry_run.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | Refund reason | |
| receipt_id | Yes | Receipt id (rcpt_…) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the annotations: it discloses that no external call is made ('Does not contact Stripe or Link'), that the budget is restored, that the operation is once-only (reinforcing idempotentHint=true), and the exact refusal condition. Annotations cover the mutation/idempotency profile, and the description adds the missing operational context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, zero filler: the action is front-loaded, followed by the human-confirmation requirement and the two key behavioral constraints. Every sentence carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive mutation with no output schema, the description covers the essentials: what it reverses, the budget side effect, the external-call behavior, the once-only semantics, and the refusal precondition. Nothing needed to call it safely is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there are only 2 parameters, so the schema already documents receipt_id and reason. The description adds no format, constraint, or meaning for either parameter beyond what is already encoded, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Reverse') and resource ('settled dry-run receipt') plus the resulting side effect ('restore that amount to the remaining budget'). This is clearly distinguishable from the sibling budget/proposal tools without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly flags that this is 'the step a human should confirm,' which tells the agent when to pause for approval, and states the precondition ('Refuses receipts that are not settled_dry_run'). It does not name an alternative tool, but none of the siblings overlaps, so this is a clear context rather than a full routing rule.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_budgetSet budgetAIdempotent
Overwrite the local spend budget. Changes the ledger limit only. Does not spend money, approve a card, settle a purchase, or contact Stripe or Link. Passing the same amount again does not add to the budget.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | Optional note for this budget | |
| currency | No | Currency code; default USD | |
| amount_usd | Yes | Budget amount in USD |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare idempotentHint=true and openWorldHint=false, so the description's job is to add nuance. It does so usefully: 'Passing the same amount again does not add to the budget' clarifies that idempotency here means overwrite rather than increment, and 'Changes the ledger limit only' scopes the blast radius. It omits any permission or auth requirements, keeping it below 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and effect, then the exclusions and the idempotency caveat. Every sentence carries distinct information with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter mutation with no output schema and annotations covering the safety profile, the description covers what the tool changes, what it does not change, and repeat-call behavior. It lacks any statement of required permissions or the resulting state after the overwrite, which keeps it from a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so amount_usd, currency, and note are already documented in the schema; the baseline of 3 applies. The description adds no format, range, or unit detail beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Overwrite the local spend budget') plus the precise effect ('Changes the ledger limit only'). It is clearly distinguishable from siblings like propose_purchase and refund_receipt, which the description explicitly disclaims.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The negation list ('Does not spend money, approve a card, settle a purchase, or contact Stripe or Link') effectively tells the agent when this tool is not the right one, which is strong routing guidance. It stops short of naming the sibling tool an agent should use instead, so it is not a full 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
7 tool updates
v0.1.1- First observed
decide_proposal - First observed
get_budget - First observed
get_receipts - First observed
list_proposals - First observed
propose_purchase - First observed
refund_receipt - First observed
set_budget
TDQS
Scored across 7 tools
Each tool maps to a distinct resource+action: budgets (set/get), proposals (propose/list/decide), and receipts (get/refund). Descriptions explicitly clarify boundaries, e.g. set_budget vs propose_purchase vs decide_proposal, leaving no real overlap.
Every tool follows a clean verb_noun snake_case pattern (set_budget, get_budget, propose_purchase, list_proposals, decide_proposal, get_receipts, refund_receipt). No mixing of conventions.
Seven tools is well-scoped for a spend-governance ledger, each covering a necessary stage of the budget/proposal/receipt lifecycle with no filler.
Covers the full lifecycle: budget set/read, proposal create/list/decide, receipt read/refund. Minor gaps like filtering or clearing stale proposals exist but core workflows have no dead ends.
Related MCP Connectors
MCP server for Boson Protocol — on-chain agentic commerce for physical & digital goods.
Marketplace MCP server for live venture-capital raises and agent-mediated dealflow.
MCP server for mandates, delegation, policy-gated execution, credential grants, and audit.
Agent-native marketplace. Bootstrap, list inventory, search, negotiate, and trade via MCP.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceMCP server for AI agent billing. Preflight spend checks before agent runs. Post-execution usage billing via two MCP tools: preflight() and record_event().7MIT
- AlicenseNot gradedqualityCmaintenanceBudget management and cost tracking MCP server for autonomous agents, enabling budget creation, cost recording, spending projections, and alert rules.MIT
- FlicenseNot gradedqualityBmaintenanceAn MCP server that enables AI agents to safely interact with a double-entry payments ledger, enforcing idempotency, policy-based access control, and human-in-the-loop approval for high-value actions.-
- AlicenseNot gradedqualityCmaintenanceMCP server that enables AI agents to request human approval before spending money, check approval status, verify signed tokens, and manage API keys.13 npmMIT