Skip to main content
Glama

proxy-shopping-web

Buy from cash-only shops with crypto, or earn as a proxy shopper: the browser app, guided demo, TypeScript protocol library and MCP server of proxy-shopping, a P2P network where a local shopper buys for remote users who pay through a 2-of-3 escrow with timelocks. Agents: MCP server io.github.pad01g/proxy-shopping (packages/mcp), skills npx skills add pad01g/proxy-shopping-go, AGENTS.md; overview for machines: https://pad01g.github.io/proxy-shopping-docs/llms.txt.

Documentation in 14 languages: https://pad01g.github.io/proxy-shopping-docs/

Browser side of proxy-shopping (spec: ../proxy-shopping-go/docs/spec.md, lab: docs/lab.md there).

packages/core   @proxy-shopping/core — protocol library (TypeScript, ESM, browsers and Node 22)
apps/web        React + Vite single-page app for every role (Japanese UI)
apps/demo       the lab's integrated demo page: every role with its own key in one browser, a guide per scenario
scripts/        docker helpers for integration and browser tests

Host tools are not required; everything runs in docker.

# install, build, unit tests
docker run --rm -v "$PWD":/src -v ps-npm:/root/.npm -w /src node:22-bookworm npm ci
docker run --rm -v "$PWD":/src -v ps-npm:/root/.npm -w /src node:22-bookworm npm run build
docker run --rm -v "$PWD":/src -v "$PWD/../proxy-shopping-go":/proxy-shopping-go:ro -v ps-npm:/root/.npm -w /src node:22-bookworm npm test

# integration: bitcoind signet + anvil (lab contracts) + nostr-rs-relay
scripts/integration.sh
# the built web app in Chromium against a mini lab (TS shopper, faucet, escrow …)
scripts/e2e-web.sh
# the demo in mock mode (GitHub Pages build) in Chromium: every scenario, nothing leaves the page
scripts/e2e-demo-mock.sh

# web image (config.json is mounted at runtime)
docker build -f apps/web/Dockerfile -t proxy-shopping-web .
docker run -p 8080:80 -v $PWD/config.json:/usr/share/nginx/html/config.json:ro proxy-shopping-web

Mounting ../proxy-shopping-go read-only lets the unit tests check docs/test-vectors.json and lab/keys/public.json. Without the vectors file npm test fails, so CI cannot pass silently; set SKIP_VECTORS=1 to run the other tests without it.

apps/demo

Try it in your browser: https://pad01g.github.io/proxy-shopping-web/ (everything simulated in the page). That build runs in mock mode (?mock=1, header switch モック / Mock): the relays, the BTC chain (with script checks), a real EVM with the lab's Safe and escrow-module bytecode, the faucet, the rates, the shops and the shopper-1 node run inside the browser (a SharedWorker shared by all windows), while the four roles use core's real clients. Details, what is simulated and its limits: apps/demo/README.md. The mock e2e (all scenarios, English, separate windows, no request leaving the page): scripts/e2e-demo-mock.sh; Pages is deployed by .github/workflows/pages-demo.yaml.

The integrated demo for the docker compose lab (../proxy-shopping-go, compose service demo, http://localhost:8888/; usage in ../proxy-shopping-go/docs/lab.md「デモ画面」). User, escrow, operator and coordinator each have their own key (localStorage, lab only) and their own Session + client in the page; the shopper is the lab's Go node. A guide walks through seven scenarios (normal BTC / USDC, dispute refund, sold out, risky shop, fraudulent escrow, T2 refund); ?role=user / ?role=escrow,operator,coordinator runs only those roles in a window. The page is in Japanese or English (header switch 日本語 / English, kept in localStorage; ?lang=en|ja overrides it): one message catalog per language in apps/demo/src/i18n (ja.ts defines the keys, en.ts must match them), core's timeline lines are rendered from their kind. Protocol logic is core's (UserClient, EscrowClient, OperatorClient, CoordinatorClient); the page wires them over core's MappedTransport (logical relay URLs such as wss://relay-1.test stay protocol-visible, the connection goes to the demo server's /relay-1).

docker build -f apps/demo/Dockerfile -t proxy-shopping-demo .   # static files; the lab mounts its nginx.conf + demo-config.json
DEMO_SERVER=http://localhost:8888 npm run dev -w @proxy-shopping/demo   # dev server proxying to a running lab demo server

Scenario steps are data (apps/demo/src/scenarios); test ids in apps/demo/TESTIDS.md; the e2e that follows the guide is ../proxy-shopping-go/e2e/src/demo (docker compose run --rm runner demo).

Related MCP server: arbitova-mcp-server

@proxy-shopping/core

Entry points: @proxy-shopping/core (portable), /node (+ FileStorage), /browser (+ IndexedDBStorage), /testing (in-memory relays/chain, scripted shopper, createWorld), /testing/node (bitcoind RPC + Esplora shim). Build output with declarations is in dist/.

area

main exports

keys

KeySet.fromMnemonic (nostr, order keys, escrow tpub, P2WPKH wallet, EVM), orderIndex, escrowPubkeyFromXpub, generateMnemonic, LocalSigner, Nip07Signer

nostr

signInner, giftWrap, unwrap, isValidInner, Messenger (k relays, ack, retry, dedupe, persisted), PoolTransport, MappedTransport, message body types, MSG

trust

effectiveCombinations, matchingEntries, covers, latestByAddress, parsers and *Template builders for 30500–30503/10050, TrustDirectory

fx

FrankfurterSource, CoingeckoSource, ChainlinkSource, StaticSource, computePair, getRate, checkQuoteRate

btc

witnessScript, p2wshAddress, EsploraClient, buildFundingTx, buildEscrowSpend, signEscrowInput, finalizeEscrowInput (multisig / T1 / T2), PSBT base64 helpers

evm

safeInitializer, predictSafeAddress, orderSafeAddress, releaseSafeTx, splitSafeTx, encodeMultiSend, safeTxHash, signSafeTx, packSignatures, EvmClient (deploy, fund, exec, module, bond)

delivery

encryptAddress, decryptAddress, sealDelivery, unwrapDeliveryKey

flows

Session, UserClient, EscrowClient, OperatorClient, CoordinatorClient, ShopperProfile, checkQuote, checkTimelock, btcPayoutProblems / safePayoutProblems (§4.10 templates), escrowSpent

checks

parseBody / BODY_SCHEMAS (every message body is schema-checked at the Messenger boundary), signKeyProof* / verifyRequestKeyProof (§4.4.1), endpointProblem, encryptWithPassphrase

Notes

  • The mnemonic is stored in the origin's IndexedDB encrypted with a passphrase (PBKDF2-SHA256, 600k iterations, AES-GCM via WebCrypto); plaintext only if the user explicitly opts out at onboarding. WebCrypto needs a secure context (https or localhost) — scripts/e2e-web.sh forwards localhost:8080 in the browser container for that.

  • Funds never move without a click and an in-page confirmation (fund, resume funding, release, countersign, refunds). The shopper's cooperative refund is shown for review, never auto-signed. Terminal states (completed, settled, refunded) wait for the chain, also after our own broadcast (BTC /tx/{txid}/outspend/{vout}; USDC: Safe balance below lock_amount and a receipt with a Transfer out of the Safe — anyone can send dust to a Safe, §4.8).

  • Funding persists every transaction before it is sent (BTC txid + raw tx; EVM: signed locally, hash + raw tx stored, then sent), so an interrupted funding is resumed with the same transactions, never paid twice.

  • The escrow opens a case only when a notice / dispute names a request whose request + quote recompute to the funded output (P2WSH, or the Safe with its owners / module configuration) on chain (§4.7); its ruling fee is exactly dispute_fee_bps of what is split, and it owes rulings only for upfront fees ≥ its own published minimum.

  • config.json fields beyond the endpoints: timelock_policy (§4.5.1; names as in proxy-shopping-go/lab/web-config.json, spec defaults when absent), max_clock_skew_seconds (§4.5.1: largest accepted difference between the chain's clock — BTC tip header time, EVM latest block time — and ours, default 7200; the lab sets 315360000 because anvil's time is warped), allow_private_endpoints (lab only: http/ws and private hosts), max_fee_rate (sat/vB, default 50). Endpoints must otherwise be https / wss; relays from peers' kind 10050 are limited to 8 public wss relays.

  • Messaging (§4.10): 120 messages / min per sender after EOSE, 60 / min for all non-counterparty senders together, higher per-sender limits for the stored backlog (older pages are read when the first 1000 wraps are full); messages no attached role accepts are neither stored nor acked. Sends go to k (2) inbox relays, the next ones only when one fails; resends 1, 2, 4, 8, … get a fresh wrap. A signed inner is at most 28000 bytes (§4.9); dispute evidence is split over several messages.

  • The Safe address is predicted with the factory's proxyCreationCode only if its keccak256 is a known Safe v1.4.1 build (the lab's and the canonical one, KNOWN_PROXY_CREATION_CODE_HASHES).

  • Only one tab runs the protocol at a time (Web Locks, BroadcastChannel fallback).

  • apps/web/nginx.conf sets CSP, X-Content-Type-Options, Referrer-Policy, Permissions-Policy, X-Frame-Options; HSTS belongs to the TLS terminator. Production builds have no source maps.

  • The browser talks to Esplora, the EVM RPC, rate sources and the faucet directly, so those endpoints must send CORS headers.

  • UI test ids: apps/web/TESTIDS.md.

Contributing — pull requests welcome

Pull requests are welcome: shop drivers for shopper-bot, new payment methods and chains, translations, protocol and security reviews, bug fixes. Want to earn in your own town? You need nobody's permission: make a coordinator key, delegate to your operator key and list yourself as a shopper for your region (guide: https://pad01g.github.io/proxy-shopping-docs/en/quickstart/, section 3); a home machine reached over Tailscale or any VPN is enough. To be found by everyone, open a pull request to https://github.com/pad01g/proxy-shopping-registry.

License

MIT (LICENSE).

Available Tools

19 tools
accept_quoteAccept a quoteA

Step 3: accept a quoted order (sends a signed acceptance to the shopper; nothing is paid yet — fund_order pays). Only for status quoted. A quote that failed validation is refused. A quote whose rate deviates strongly (> 10 %) from our sources, or could not be checked, is not accepted until you call again with acknowledge_rate_deviation: true — ask your human first.

ParametersJSON Schema
NameRequiredDescriptionDefault
order_idYesorder id (32 hex characters) from request_quote or list_orders
acknowledge_rate_deviationNotrue to accept although the quote needs an acknowledgement (strong rate deviation or rate not checkable)

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare the generic mutation profile (readOnlyHint=false, idempotentHint=false, destructiveHint=false). The description adds meaningful behavior beyond that: what the call actually sends, that no funds move at this step, the validation-refusal rule, and the rate-deviation acknowledgement gate with its >10% threshold.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the step number and the highest-value clarification (nothing is paid yet), followed by eligibility and the gating condition. No filler and no repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet nothing an agent needs to invoke this correctly is missing: pipeline position (step 3), state prerequisite, failure behavior, the acknowledgement flag path, and the human-escalation requirement are all stated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the description's explanation of acknowledge_rate_deviation ('strong rate deviation or could not be checked') largely restates the schema text. The >10% threshold is the only genuinely additive detail, so this sits at the baseline the schema already earns.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('accept a quoted order') and immediately clarifies the material consequence: a signed acceptance is sent but nothing is paid yet, with fund_order named as the tool that pays. This distinguishes it cleanly from siblings like request_quote, fund_order, and cancel_order.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit preconditions ('Only for status quoted'), explicit failure routing ('A quote that failed validation is refused'), and an explicit escalation rule ('ask your human first' before acknowledge_rate_deviation). When-to-use, when-it-fails, and who-to-ask are all covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

accept_refund_offerAccept the shopper's refundA
Destructive

Moves money: co-sign and broadcast a cooperative refund the shopper offered (e.g. the item was sold out or delivery failed), returning the funds to your wallet. Refused if the offer does not pay you as the refund template requires. Without confirm: true it shows the offer and its check. Errors if there is no offer (get_order shows refund_offer).

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNotrue to really do it (sign / send / broadcast); omitted or false = preview only, nothing happens
order_idYesorder id (32 hex characters) from request_quote or list_orders

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (destructive, non-idempotent, open-world), and the description adds behavior annotations cannot: a preview mode that shows the offer and its check, a refusal condition tied to the refund template, and an error condition when no offer exists. These are exactly the preconditions an agent needs before a money-moving call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with 'Moves money:' to signal stakes, then two dense sentences covering effect, failure mode, preview gating, and prerequisite. Parenthetical examples add length but carry useful context rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter destructive tool with no output schema, the description covers effect, confirmation gating, refusal semantics, error routing, and the lookup needed to find an offer. Nothing essential for a correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description still adds value by specifying what the non-confirm path actually returns ('shows the offer and its check'), which the schema only summarizes as 'preview only, nothing happens'. order_id semantics are left entirely to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: co-sign and broadcast a cooperative refund the shopper offered, with the effect spelled out (funds return to your wallet). The word 'cooperative' plus 'the shopper offered' distinguishes it from the unilateral sibling refund_after_timelock.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete trigger scenarios (item sold out, delivery failed) and a prerequisite pointer to get_order, which shows refund_offer when no offer exists. It does not explicitly contrast with refund_after_timelock or accept_quote, so routing is clear but not fully disambiguated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

become_shopperEarn as a proxy shopperA
Read-onlyIdempotent

Return a step-by-step plan to earn fees as a proxy shopper (buying at local shops for remote users who pay into the escrow): requirements (always-online host, Docker, the Go node and shopper-bot images, card or cash regions), fees and risks, an honest status of the network, a ps-main node config and compose file, and how to get listed through the registry (then registry_entry). Local, no network, nothing is started.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoyour shopper display name
fee_bpsNoyour fee in basis points (100 = 1 %)
regionsNoregions you serve, as region codes (e.g. JP-13)
cash_regionsNoregions where you can pay cash in person (e.g. JP-13-13104)

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and openWorldHint=false, and the description reinforces rather than contradicts them while adding genuinely new context: 'Local, no network, nothing is started.' It could say more about the plan's size or format, but the safety/idempotency profile is well-covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence, front-loaded with the purpose before the content list, and the trailing 'Local, no network, nothing is started' is a useful scoping qualifier. The embedded list is dense but every item earns its place by telling the agent what the plan will contain.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, side-effect-free plan generator with no output schema, the description tells the agent the plan's contents, that nothing is executed, and what the follow-up step is. Nothing material is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters (name, fee_bps, regions, cash_regions) are already documented in the schema. The description only lightly gestures at them ('card or cash regions', 'fees'), adding no new syntax, format, or constraint detail. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Return a step-by-step plan to earn fees as a proxy shopper') and then enumerates exactly what the plan contains (requirements, fees/risks, status, config file, listing path). This clearly distinguishes it from action siblings like fund_order or accept_quote, and it names registry_entry as the natural follow-up.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives context ('Local, no network, nothing is started') and routes the agent onward with 'how to get listed through the registry (then registry_entry)'. It doesn't state explicit when-not-to-use conditions against other informational siblings like network_info, so it falls short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_orderCancel an order before fundingA

Cancel an order that is not funded yet (requested, quoted or accepted) and tell the shopper (best effort). Refused once funding has started — then use open_dispute, accept_refund_offer or refund_after_timelock. Nothing is paid or refunded.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonNoreason sent to the shopper (default "cancelled by user")
order_idYesorder id (32 hex characters) from request_quote or list_orders

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare it is not read-only and not destructive; the description adds genuinely new behavior: the state-machine precondition, the refusal boundary, a best-effort shopper notification side effect, and the guarantee that no money moves. It stops short of stating idempotency or permission requirements, so 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no filler, front-loading the eligibility constraint and the failure/alternative path. Every clause carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter mutation with no output schema, the description supplies the precondition, the refusal case, the fallback tools, and the absence of financial side effects. Nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both order_id (32-hex format, source tools) and reason (max length, default) are already fully documented. The description's only added signal is the implied 'best effort' delivery of the reason to the shopper, which is marginal; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (cancel) and resource (order) and narrows the scope to unfunded orders, enumerating the eligible states (requested, quoted, accepted). An agent can immediately tell this apart from fund_order, refund_after_timelock, or open_dispute without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly gives the use condition (order not funded yet), the exclusion (refused once funding has started), and names three concrete alternative tools for the funded case. This is a complete routing rule.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

confirm_receiptConfirm receipt and pay the shopperA
Destructive

Step 5, moves money: the items arrived and are right, so sign the escrow payout to the shopper (release); the shopper co-signs and broadcasts it. Irreversible. Without confirm: true it only returns the payout (amount, recipient) and a warning if the shopper has not reported delivery. With confirm: true it signs and waits up to wait_seconds for the payout on chain. If the items did not arrive or are wrong, use open_dispute instead.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNotrue to really do it (sign / send / broadcast); omitted or false = preview only, nothing happens
order_idYesorder id (32 hex characters) from request_quote or list_orders
wait_secondsNohow long to wait for the on-chain completion, in seconds (default 20)

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructive=true, readOnly=false, idempotent=false, but the description adds materially beyond them: irreversible, the two-phase preview/confirm flow, the shopper co-sign and broadcast, a warning when the shopper has not reported delivery, and the wait_seconds polling window. That is genuine behavioral detail an agent needs before committing money.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the riskiest fact ('moves money', 'Irreversible') and the preview/commit distinction before the alternative-tool routing. Every sentence carries decision-relevant information; nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, and the description compensates by describing what the preview returns and what happens on confirm. Combined with annotations covering the safety profile, an agent has everything needed to invoke this correctly or route elsewhere.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline would be 3, but the description adds semantics the schema does not: what confirm:true does to the flow (signs and waits up to wait_seconds on chain) and what the preview returns (payout amount, recipient, possible warning). It goes modestly beyond the structured fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: sign the escrow payout to release funds to the shopper. It also positions itself in the workflow ('Step 5, moves money') and contrasts with a named sibling, open_dispute, so an agent can distinguish it without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('the items arrived and are right'), an explicit when-not with the alternative ('if the items did not arrive or are wrong, use open_dispute instead'), and the preview-vs-commit condition for confirm. Nothing about selection is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

countersign_rulingCountersign the rulingA
Destructive

Moves money: add your signature to the escrow's ruling transaction (2 of 3) and broadcast it, settling the dispute. Refused if the transaction does not pay the ruled split. Without confirm: true it returns the same review as review_ruling. Ask your human before confirming.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNotrue to really do it (sign / send / broadcast); omitted or false = preview only, nothing happens
order_idYesorder id (32 hex characters) from request_quote or list_orders

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, idempotentHint=false, and readOnlyHint=false. The description adds real value beyond them: it specifies that money moves and the dispute is settled, the refusal condition, and the confirm-gated preview behavior — context an agent cannot get from the annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the critical fact ('Moves money') and packs the precondition, preview behavior, and safety directive into a compact block. Dense but every clause earns its place; slightly clausal but no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, but the description explains the return behavior (returns the same review as review_ruling when not confirmed) and the main failure mode. Combined with destructive annotations, an agent has enough to invoke it safely; only exhaustiveness of error cases is unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already documents both confirm (preview vs real) and order_id (32-hex from request_quote/list_orders). The description reinforces the confirm semantics but adds little beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (countersign/sign+broadcast) and resource (the escrow's ruling transaction), plus the concrete effect — 'Moves money' and 'settling the dispute'. It implicitly distinguishes itself from review_ruling by naming it as the fallback behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit preconditions ('Refused if the transaction does not pay the ruled split'), the preview-vs-execute condition ('Without confirm: true it returns the same review as review_ruling'), and a human-in-the-loop directive ('Ask your human before confirming'). This is close to a full when/when-not/guardrail set.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

export_backupExport the recovery wordsA
Read-onlyIdempotent

Show this agent's 12-word BIP39 mnemonic, which controls the identity and every order's escrow key, with restore instructions. Only with confirm: true — the words then stay in the conversation, and anyone who sees them can take the funds; use only when your human asks for a backup. Local, no network.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNotrue to really do it (sign / send / broadcast); omitted or false = preview only, nothing happens

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Discloses critical consequences beyond the annotations: the words persist in the conversation, anyone who sees them can take the funds, and the operation is local with no network. readOnlyHint/idempotentHint correctly reflect that nothing is mutated, so there is no contradiction and the risk disclosure adds real value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with what is exported and what it controls, followed by the gating condition and the risk. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-output-schema secret-export tool, the description covers the risk, the confirm gate, and network locality. It leaves the preview mode (confirm omitted/false) unmentioned, a minor gap for an agent deciding whether a bare call is safe.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the confirm boolean and its semantics are already documented in the schema. The description reinforces the confirm: true requirement but omits the schema's 'omitted or false = preview only' behavior, so it adds little beyond the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Show this agent's 12-word BIP39 mnemonic') and immediately scopes what it controls (identity and every order's escrow key). This is clearly distinguishable from siblings like wallet or registry_entry without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear selection condition ('use only when your human asks for a backup') and a hard gate ('Only with confirm: true'). It does not name an alternative tool for non-backup identity inspection, so it is clear context rather than full when/when-not routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

find_offersFind proxy shoppers for a shopA
Read-onlyIdempotent

Step 1 of buying from a cash-only or crypto-unsupported shop: list the trusted proxy shopper × escrow combinations that serve this shop, region and payment method. Each offer has an index, the shopper's fee, delivery days, cash regions and order limit, the escrow's fees and dispute SLA, and which operator list (under which coordinator) vouches for it. Pass the index to request_quote with the same shop_url, region and payment. Read-only; an empty list means nobody serves that shop/region yet (network_info shows what exists).

ParametersJSON Schema
NameRequiredDescriptionDefault
regionYesthe shop's region code, prefix-matched: country 'JP', prefecture 'JP-13', Japanese municipality 'JP-13-13104'
paymentNohow you pay: btc-signet (default; the only one on ps-main) or usdc-evm
shop_urlYesthe shop's URL, e.g. https://shop.example/

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnlyHint, openWorldHint and idempotentHint, so 'Read-only' is redundant. However the description adds real behavioral context: an empty list has a specific meaning (nobody serves that shop/region yet), and it enumerates what each offer carries (fee, delivery days, cash regions, order limit, escrow fees, dispute SLA, vouching operator list) in the absence of an output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the essential 'Step 1 of...' framing and keeps the reader on a clear workflow path. It is dense and slightly long due to the offer-field enumeration, but each clause carries useful information rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description takes on the burden of describing return contents (offer index, fees, delivery days, escrow SLA, operator list) and explains how to consume the index. Combined with the sibling routing, an agent has everything needed to call it and act on the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents region prefix-matching, the payment enum and shop_url format. The description reinforces the coupling of shop_url/region/payment across find_offers and request_quote but adds no new per-parameter syntax beyond what the schema provides, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: list proxy shopper × escrow combinations serving a given shop, region and payment method. This is clearly distinct from siblings like network_info (what exists) and request_quote (the next step), which the description itself names.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames itself as 'Step 1 of buying from a cash-only or crypto-unsupported shop', names the follow-up tool (request_quote) and what to pass to it, and routes the no-results case to network_info. When-to-use, next-step and fallback are all covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fund_orderPay into the escrowA
Destructive

Step 4, moves money: pay the quoted lock amount into the per-order 2-of-3 escrow (BTC P2WSH address or USDC Safe) plus the escrow upfront fee, from this wallet. Only for status accepted/funding. Without confirm: true it only returns the recipients, amounts and network fee (a preview; call it first). With confirm: true it signs and broadcasts. The funds can then leave only with 2 of 3 signatures, by the shopper alone after T1, or by you alone after T2. Ask your human before confirming.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNotrue to really do it (sign / send / broadcast); omitted or false = preview only, nothing happens
order_idYesorder id (32 hex characters) from request_quote or list_orders

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructive/non-idempotent/openWorld, and the description goes well beyond them by disclosing what the call does in each mode (preview returns recipients, amounts, network fee; confirm signs and broadcasts), the custody conditions under which funds can leave (2-of-3, shopper after T1, agent after T2), and the approval requirement. That is substantive behavioral context an agent cannot infer from the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action ('Step 4, moves money: pay the quoted lock amount...') before the conditions and modes. It is dense with several clauses per sentence, but each clause carries operational or safety information rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent money-moving tool with no output schema, the description supplies the state precondition, the preview/execute duality, the approximate return contents, the custody/timelock release rules, and the human-approval requirement. Nothing material to calling it safely is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds workflow meaning the schema does not: confirm=false is a preview that should be called first, and confirm=true signs and broadcasts. order_id is only referenced indirectly via 'this wallet'/'per-order', so it is not fully compensated, but the two-phase semantics of confirm are clearly enriched.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: paying the quoted lock amount into the per-order 2-of-3 escrow (BTC P2WSH or USDC Safe) plus the escrow fee. It is unmistakable against siblings like request_quote, accept_quote, or confirm_receipt, and even labels itself as step 4 in the flow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit preconditions ('Only for status accepted/funding'), an explicit ordering rule for the preview mode ('Without confirm: true ... call it first'), and an explicit human-in-the-loop gate ('Ask your human before confirming'). Both when-to-use and how-to-sequence are covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_orderOrder statusA
Read-onlyIdempotent

Show one order: status, quote and its validation, funding, purchase, tracking, dispute, ruling, refund offer and settlement, the recent timeline, and next_steps telling which tool to call next. Optionally waits (bounded) until the order reaches one of wait_for, e.g. ["quoted","rejected"] after request_quote. Read-only (local order state, kept up to date by the running session).

ParametersJSON Schema
NameRequiredDescriptionDefault
order_idYesorder id (32 hex characters) from request_quote or list_orders
wait_forNoreturn as soon as the status is one of these (e.g. ["quoted","rejected"] or ["completed"])
wait_secondsNoupper bound for wait_for, in seconds (default 30)

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint/idempotentHint/openWorldHint=false, so safety is covered. The description adds genuine context beyond that: the optional bounded wait until a target status is reached, and that state is local and kept current by the running session — useful behavior an agent cannot get from annotations alone.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the returned contents before the waiting behavior in two dense sentences with no filler. It is somewhat run-on in the first sentence, but each clause (returned sections, next_steps, bounded wait) earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully enumerates the returned sections and the next_steps routing hint, and it explains the wait semantics. Pitch and pagination details are absent but not critical for a single-order read, so coverage is close to complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the schema already documents order_id origin, the wait_for enum, and wait_seconds bounds/default. The description largely restates the wait_for example already present in the schema, so it adds little beyond the baseline expected when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (show) and resource (one order) and enumerates exactly what is returned — status, quote/validation, funding, purchase, tracking, dispute, ruling, refund, settlement, timeline, next_steps. The singular 'one order' plus the sibling list_orders makes the scope unambiguous without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: use it to inspect a single order and get next_steps telling which tool to call next, and it explains the wait_for case ('after request_quote'). It does not explicitly name when NOT to use it versus list_orders or a more targeted sibling, but the intended usage is easy to infer.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_ordersList ordersA
Read-onlyIdempotent

List this agent's orders (from its data dir), newest first, with status, shop, items, lock amount and next steps. Use it to find an order id or to see what needs attention; get_order shows one order in full. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNomost orders to return (default 20)
statusNoonly orders in these statuses

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent and closed-world, and the description restates read-only while adding useful behavior: source is the agent's data dir and results come newest first. It does not discuss pagination or result-size limits beyond what the schema already documents.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One tight sentence for purpose/scope/returned fields, one for usage and the alternative, and a two-word safety tag. No wasted text, front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description enumerates the returned fields (status, shop, items, lock amount, next steps), covers usage and the sibling, and annotations cover the safety profile. Nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema documents both parameters at 100% coverage (limit default, status enum list), so it carries the burden. The description's mention of 'status' refers to a returned field rather than the filter, adding no syntax or meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (this agent's orders) plus scope, ordering ('newest first') and the fields returned. It is clearly distinguishable from get_order, which is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use cases ('find an order id', 'see what needs attention') and names the alternative tool (get_order) with the condition that selects it (one order in full).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

network_infoNetwork statusA
Read-onlyIdempotent

Show which proxy-shopping network this server is on (ps-main = public Nostr relays + BTC signet, or the local lab), its relays, trusted coordinators, chains, timelock policy, and how many shopper × escrow combinations are trusted right now (with their regions and operator lists). Call it first in a session: on a new network there may be no shoppers yet, and then find_offers will return nothing. Read-only; re-reads the trust lists from the relays and the registry unless refresh is false. Returns a summary and the details as JSON.

ParametersJSON Schema
NameRequiredDescriptionDefault
refreshNore-read the trust lists from the relays and the registry (default true); false uses the last snapshot and is faster

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and openWorldHint, so the safety profile is covered. The description adds real behavioral context on top: re-reads trust lists from relays/registry unless refresh is false, and returns a summary plus JSON details. It does not cover failure modes or latency, but the added refresh semantics are valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose before the dense enumeration of reported fields, and every sentence carries information (purpose, contents, routing advice, refresh behavior, return shape). It is dense with domain jargon and longer than needed, which keeps it out of the top tier.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description must characterize returns—and it does ('Returns a summary and the details as JSON'). Combined with the session-first guidance and refresh semantics, an agent has enough to call it correctly. Some network-state detail (e.g., regions/operators) is described but not elaborated, leaving minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the refresh parameter's meaning is already fully documented in the schema. The description restates refresh indirectly via 'unless refresh is false' but adds no syntax or default detail beyond what the schema provides. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Show) and a concrete resource (which proxy-shopping network this server is on), then enumerates exactly what is reported: relays, coordinators, chains, timelock policy, and trusted shopper × escrow combinations. It also distinguishes itself from the sibling find_offers by explaining the dependency relationship between them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear directive: 'Call it first in a session,' and explains the consequence of skipping it (on a new network find_offers returns nothing). This is strong when-to-use guidance tied to an outcome. It stops short of explicit exclusions or naming alternatives for adjacent cases, so it falls just below the top tier.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_disputeOpen a disputeA

Ask the escrow to decide a funded order: items not delivered, wrong item, or the shopper not releasing. Sends the claim with all signed order messages as evidence (copy to the shopper) and the key that lets the escrow decrypt the delivery address. No funds move now; the escrow later rules a split, which you check with review_ruling and execute with countersign_ruling. Without confirm: true it only shows what would be sent.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesyour explanation for the escrow (facts, dates, tracking)
claimYeswhat went wrong
confirmNotrue to really do it (sign / send / broadcast); omitted or false = preview only, nothing happens
order_idYesorder id (32 hex characters) from request_quote or list_orders
requested_splitNothe split you ask for, as integers in sats (BTC) or USDC base units (6 decimals)

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only give the generic mutation profile (readOnly=false, idempotent=false, destructive=false). The description adds genuinely non-obvious behavior: no funds move now, the claim ships signed order messages copied to the shopper, and a decryption key for the delivery address is transmitted, plus that omitting confirm only previews.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed sentences, front-loaded with the core action and consequences before the workflow pointer. Every clause carries information; nothing is redundant filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-param, nested, non-idempotent mutation with no output schema, the description covers purpose, evidence payload, side-effect timing, follow-on calls, and the preview gate. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all five parameters (including the claim enum, requested_split shape, and confirm semantics) are already documented in the schema. The description reinforces the confirm behavior but adds no format or constraint detail beyond it, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('ask the escrow to decide a funded order') and enumerates the triggering conditions (not delivered, wrong item, not released). It also names the downstream siblings (review_ruling, countersign_ruling) so an agent can place it in the dispute workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use conditions and the follow-on sequence ('the escrow later rules a split, which you check with review_ruling and execute with countersign_ruling'). It also clarifies the preview-vs-real toggle, which governs whether it should be invoked at all.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

refund_after_timelockTake the funds back after T2A
Destructive

Moves money, last resort: after the T2 timelock you alone can take the locked funds back to your wallet (e.g. the shopper disappeared). Without confirm: true it shows T2, the current block height or chain time, whether T2 is reached, and the amount. Before T2 the chain rejects it; prefer accept_refund_offer or open_dispute while the shopper responds.

ParametersJSON Schema
NameRequiredDescriptionDefault
confirmNotrue to really do it (sign / send / broadcast); omitted or false = preview only, nothing happens
order_idYesorder id (32 hex characters) from request_quote or list_orders

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true and readOnlyHint=false, but the description adds real context: the two-mode preview/execute behavior, what the preview returns (T2, current block height or chain time, whether T2 is reached, amount), and the hard failure condition before T2. It stops short of stating permission/irreversibility details beyond the timelock gating, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and its 'last resort' framing, then the preview semantics, then the failure/alternative routing. Three dense sentences with no filler; every clause carries decision-relevant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by describing what the preview returns and when the operation is rejected. With full schema coverage, annotations, and these additions, an agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (confirm, order_id) are already fully documented in the schema, including the preview-vs-execute meaning of confirm. The description restates that preview behavior rather than adding new parameter-level detail, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action (take locked funds back to your wallet), the trigger (after the T2 timelock), and the actor (you alone), naming the scenario it exists for (shopper disappeared). It is clearly distinguishable from siblings like accept_refund_offer and open_dispute, which it explicitly points at.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('last resort', 'after the T2 timelock'), when-not ('before T2 the chain rejects it'), and named alternatives ('prefer accept_refund_offer or open_dispute while the shopper responds'). An agent has a complete decision rule without reading any schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

registry_entryRegistry entry for getting listedA
Read-onlyIdempotent

Return the exact JSON file and path to add in a pull request to github.com/pad01g/proxy-shopping-registry, so a shopper, escrow, operator or coordinator gets listed once the maintainer merges it: shoppers/.json {pk, contact, description, regions (cash regions), payments, escrows}; escrows/.json {pk, contact, description, sla_days}; operators/.json {pk, contact, description, regions}; coordinators/.json {pk, contact, description, url?, bundle?}. Uses this data dir's identity pubkey unless pubkey is given (a shopper entry must carry the shopper node's key). Local; opens no pull request itself.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNocoordinator: https URL of your page
nameYesfile name: a-z, 0-9 and - (other characters are turned into -)
roleYeswhich list to join
bundleNocoordinator: https URL of your signed events.json
pubkeyNo64 hex Nostr public key (psctl keys: nostr_pubkey)
contactYeshow reviewers and users reach you, e.g. "github:<user>" or "nostr:npub1…"
escrowsNoshopper (required): names of the escrows/<name>.json you work with
regionsNoshopper: your cash regions; operator: where you list (JP, JP-13, JP-13-13104)
paymentsNoshopper: default ["btc-signet"] (the only payment on ps-main)
sla_daysNoescrow: most days from a dispute to your ruling (default 14)
descriptionYesone or two sentences: what you do, where, how

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint and openWorldHint=false, and the description is consistent with them. Beyond the annotations it adds genuine behavior: no pull request is opened, the local data dir identity pubkey is used by default, and a shopper entry must carry the shopper node's key. It stops short of describing validation or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence delivers purpose, then a compact per-role field enumeration, then the pubkey default and the 'opens no pull request' caveat. It is dense and the inline JSON shapes are hard to scan, but nearly every clause carries non-redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must explain what comes back, and it does: the exact file and path to commit. Combined with the per-role shapes, the local-only boundary and the pubkey rule, an agent has enough to call it correctly; only error/validation behavior is unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaning the schema does not: which fields each role's file must contain (e.g. shopper requires escrows and payments, escrow requires sla_days, operators use regions) and the pubkey default rule. That per-role requirement mapping is the main value beyond the flat property list.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and deliverable: return 'the exact JSON file and path to add in a pull request' to a named registry repo. It also frames the outcome (the entity gets listed once the maintainer merges), which no sibling tool does, so an agent can distinguish it from become_shopper or report without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is clear — use this when a shopper, escrow, operator or coordinator needs to be added to the registry — and it explicitly bounds the action with 'Local; opens no pull request itself.' It does not, however, name an alternative (e.g. become_shopper) or state when not to use it, which matches the rubric's 'clear context, no exclusions' level.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reportReport a shopper or escrowA

Report the order's shopper or escrow to the operator that listed them, attaching the order's signed messages as evidence (e.g. a dishonest ruling, a quote that failed validation, no delivery). Sends one signed message; no funds move and the order is not changed otherwise. The operator may remove them from its list.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYeswhat happened
subjectYeswhom to report
order_idYesorder id (32 hex characters) from request_quote or list_orders

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false and destructiveHint=false, so the description is not carrying the whole load. It still adds real value: it explains that one signed message is sent, that no funds move and the order is otherwise unchanged, and that the operator may remove the reported party.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the action and target before the evidence and side-effect details. No filler, though the parenthetical examples add some length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description usefully covers outcome effects (no funds move, order unchanged, operator may delist) for a moderate-complexity action. It omits only repeat-call behavior and failure handling for an explicitly non-idempotent tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with the order_id format sourced to request_quote/list_orders already documented in the schema. The description only reinforces that the order's signed messages are attached as evidence; it adds no new per-parameter semantics beyond the schema baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (report), a specific target (the order's shopper or escrow), and the recipient (the operator that listed them). This is distinguishable from siblings like open_dispute or review_ruling without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The evidence examples ('a dishonest ruling, a quote that failed validation, no delivery') imply when this tool applies, but no alternative is named and no when-not condition is given. An agent must infer that this is the escalation path distinct from open_dispute.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

request_quoteRequest a quote from a proxy shopperA

Step 2: create an order with one shopper × escrow combination (offer_index from the last find_offers with the same shop_url, region and payment, or both shopper and escrow pubkeys) and wait up to wait_seconds for the quote. Sends a signed, encrypted order request over Nostr; the delivery address is encrypted for the shopper (the escrow can read it only in a dispute). Nothing is paid. Returns the order id, the price breakdown, and our validation of the quote: rate deviation from our own rate sources, the recomputed 2-of-3 escrow address, the timelocks T1/T2. If no quote arrives in time the order stays open; follow it with get_order. Next: accept_quote or cancel_order.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemsYesitems to buy at the shop (1–20 lines)
escrowNoescrow's pubkey (64 hex), instead of offer_index; needs shopper too
regionYesthe shop's region code, prefix-matched: country 'JP', prefecture 'JP-13', Japanese municipality 'JP-13-13104'
addressYesdelivery address; encrypted end to end for the shopper
paymentNobtc-signet (default) or usdc-evm; must match the find_offers call
shopperNoshopper's pubkey (64 hex), instead of offer_index; needs escrow too
shop_urlYesthe shop's URL, e.g. https://shop.example/
offer_indexNoindex of the offer in the last find_offers result (same shop_url, region, payment)
wait_secondsNohow long to wait for the quote, in seconds (default 60; 0 = do not wait)

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare the mutation/safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), which is typical for a create operation, and the description does not contradict them. It adds real behavioral value: signed/encrypted Nostr order request, address encrypted end-to-end with escrow able to read only in a dispute, 'nothing is paid', and the timeout behavior (order stays open).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and the key selection rule, then the behavioral disclosure, then the return summary and next steps. Dense but every clause carries information; slightly long, though no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still covers what is returned (order id, price breakdown, rate deviation, recomputed 2-of-3 escrow address, timelocks T1/T2) and what happens on timeout, plus the follow-up tool. Nothing material is missing for a 9-parameter flow tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds cross-call semantics the schema cannot: offer_index must reference the last find_offers result with the same shop_url/region/payment, and shopper+escrow are the alternative to offer_index. It does not restate field formats, which the schema already handles.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('create an order', 'request a quote') with scope ('one shopper × escrow combination') and explicitly frames itself as 'Step 2' of the flow. It distinguishes itself from find_offers, accept_quote and cancel_order by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing: offer_index must come from the last find_offers with matching shop_url, region and payment, or supply both shopper and escrow pubkeys instead. It names the next alternatives (accept_quote or cancel_order) and the fallback (get_order) if no quote arrives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_rulingReview the escrow's rulingA
Read-onlyIdempotent

Show the escrow's ruling for a disputed order — the split between you, the shopper and the escrow fee, and its reason — and whether the escrow's transaction really pays exactly that split. Read-only. If it matches, countersign_ruling executes it; if not, do not countersign and consider report.

ParametersJSON Schema
NameRequiredDescriptionDefault
order_idYesorder id (32 hex characters) from request_quote or list_orders

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and idempotentHint=true, so 'Read-only.' is largely redundant. The genuinely additive behavior is that this tool cross-checks the escrow's on-chain transaction against the declared split rather than merely displaying it, but that is stated implicitly through the matching logic rather than spelled out as a verification step or its failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first clause and the decision rule follows immediately. The standalone 'Read-only.' sentence duplicates the annotation and is the one element that does not earn its place, but the rest is dense and waste-free.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description carries the return-value burden and does so: the split breakdown, the reason, and the match verdict. Combined with annotations covering safety and idempotency, an agent has everything needed to call and act on the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single order_id parameter is documented in the schema with its 32-hex format and provenance (request_quote or list_orders). The description adds no format, sourcing, or constraint detail beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Show the escrow's ruling for a disputed order') and enumerates exactly what the result contains: the split across the three parties, the reason, and whether the escrow transaction actually pays that split. This is clearly distinguishable from countersign_ruling (executes) and report (escalates).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit branching guidance: if the ruling matches the transaction, countersign_ruling executes it; if it does not, do not countersign and consider report. Names both alternatives and the condition that selects each, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

walletWallet and identityA
Read-onlyIdempotent

Show this agent's identity (Nostr pubkey), its BTC signet address and balance, and where the network has USDC its EVM address with USDC and ETH balances. Use it before fund_order to check that the wallet can pay the quote, and to get the address to fund (signet faucet; on the lab: lab_faucet). Read-only: queries the chain, never signs. The key is created on first run in the data dir; export_backup shows it.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Adds context beyond annotations: 'Read-only: queries the chain, never signs' reinforces readOnlyHint and clarifies no transaction signing occurs. 'The key is created on first run in the data dir' is important side-effect/lifecycle info not in annotations, and the pointer to export_backup is useful. Could have said more about chain latency or rate limits, but this is solid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the main purpose and usage. However, the second sentence contains a grammatical error ('where the network has USDC its EVM address') that hurts readability, and the key-creation detail is buried at the end. Still reasonably compact.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param read-only tool with no output schema, the description covers what is returned, when to call it, safety, and first-run key behavior. One awkward sentence reduces polish, but an agent has enough to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

No parameters, so baseline is 4. The description complements by specifying exactly what identity/balances are returned, which is valuable since there is no output schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resources: shows identity (Nostr pubkey), BTC signet address and balance, and EVM address with USDC/ETH balances. Distinguishes itself from siblings like network_info and export_backup through concrete output contents. Slightly impeded by an awkward grammatical slip ('where the network has USDC its EVM address').

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use: 'Use it before fund_order to check that the wallet can pay the quote, and to get the address to fund'. Names the alternative funding paths (signet faucet; on the lab: lab_faucet). Routes the agent clearly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 19 tool updatesv0.1.0
    • First observedaccept_quote
    • First observedaccept_refund_offer
    • First observedbecome_shopper
    • First observedcancel_order
    • First observedconfirm_receipt
    • First observedcountersign_ruling
    • First observedexport_backup
    • First observedfind_offers
    • First observedfund_order
    • First observedget_order
    • First observedlist_orders
    • First observednetwork_info
    • First observedopen_dispute
    • First observedrefund_after_timelock
    • First observedregistry_entry
    • First observedreport
    • First observedrequest_quote
    • First observedreview_ruling
    • First observedwallet

TDQS

A4.1/5.0

Scored across 19 tools

Disambiguation4/5

Tools are nicely differentiated by explicit multi-step labeling (find_offers → request_quote → accept_quote → fund_order → confirm_receipt) and money-moving vs read-only distinctions. Some overlap clusters exist (report vs open_dispute, and the four refund/dispute settlement tools), but descriptions consistently clarify when to use each.

Naming Consistency4/5

Dominated by clear verb_noun snake_case (find_offers, request_quote, accept_quote, fund_order, confirm_receipt, open_dispute, review_ruling, cancel_order). A few deviations break the pattern: bare noun phrases network_info, wallet, registry_entry, and a lone bare verb report.

Tool Count4/5

19 tools is on the heavier side but each maps to a distinct step in a genuinely multi-step, money-handling escrow workflow where separation aids safety. Some consolidation (info/registry helpers) might be possible, but nothing feels gratuitous.

Completeness4/5

The buy → fund → receive → dispute/refund/report → settle lifecycle is thoroughly covered, including edge cases (timelock refund, cooperative refund offer, ruling verification) and seller-side onboarding. Minor gap: no shop-discovery tool, so agents must already know shop_url.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers